Workshop am 01.-02.10.2026 in Bochum

Workshop am 01.-02.10.2026 in Bochum

Eingeladene Sprecher:

Vorträge:

  • Daniel Bartl: On the role of Rademacher complexities in statistical learning
    We study the problem of learning with respect to the squared loss over a convex class of functions. It has long been believed that the sample complexity in this setting is governed by localized Rademacher complexities. We show that, assuming access to coarse information on the covariance structure of the model class, the sample complexity is instead controlled by a localized complexity associated with the limiting Gaussian process. In heavy-tailed regimes, this quantity can be significantly smaller than the Rademacher complexity.
  • Gilles Blanchard: Transductive Conformal Inference for Full Ranking
    We introduce a method based on Conformal Prediction (CP) to quantify the uncertainty of full ranking algorithms. We focus on a specific scenario where n+m items are to be ranked by some black box“ algorithm. It is assumed that the relative (ground truth) ranking of n of them is known. The objective is then to quantify the error made by the algorithm on the ranks of the m new items among the total (n+m). In such a setting, the true ranks of the n original items in the total (n+m) depend on the (unknown) true ranks of the m new ones. Consequently, we have no direct access to a calibration set to apply a classical CP method. To address this challenge, we propose to construct distribution-free bounds of the unknown conformity scores using recent results on the distribution of conformal p-values. Using these scores upper bounds, we provide valid prediction sets for the rank of any item. We also control the false coverage proportion, a crucial quantity when dealing with multiple prediction sets. Finally, we empirically show on both synthetic and real data the efficiency of our CP method for state-of-the-art ranking algorithms such as RankNet or LambdaMart.
  • Fang Han: Generative modeling for the bootstrap
    Generative modeling builds on and substantially extends the classical idea of generating synthetic data from observed samples. In this talk, I will show that this principle provides a natural and theoretically well-founded foundation for bootstrap inference. The resulting method yields statistically valid confidence intervals for both regular and irregular estimators, including settings in which Efron’s bootstrap fails. From this perspective, the generative-modeling bootstrap could be viewed as a modern extension of the smoothed bootstrap: it has the potential to mitigate the curse of dimensionality and remain effective in challenging regimes where estimators may not be root-n consistent or admit a Gaussian limit.
  • Siegfried Hörmann: Quantifying and testing dependence to categorical variables
    We suggest a dependence coefficient between a categorical variable and some general variable taking values in a metric space. We derive important theoretical properties and study the large sample behaviour of our suggested estimator. Moreover, we develop an independence test which has an asymptotic -distribution under and which is consistent against any violation of independence. The test is also applicable to the classical -sample problem with possibly high- or infinite-dimensional distributions. We discuss some extensions, including a variant of the coefficient for measuring conditional dependence.
  • Jeong Min Jeon: Fréchet partially linear additive regression for random objects
    Fréchet regression is a statistical framework for regressing random objects that take values in a general metric space. While local Fréchet regression successfully extends classical kernel regression to this setting, it suffers from the curse of dimensionality and is not well suited to mixed-covariate settings involving both discrete and continuous covariates. In this talk, we introduce Fréchet partially linear additive models (PLAMs) for metric-space-valued responses. Despite the absence of vector space operations in general metric spaces, such as addition and scalar multiplication, we successfully incorporate semiparametric modeling into the Fréchet regression framework. We establish uniform convergence rates for the proposed estimators and demonstrate their practical performance through simulation studies and a real data application.
  • Sofia Charlotta Olhede: Modelling multiplex networks This talk will discuss how to model and make inferences of multiplex networks. One point-of-view starts from a special class of models based on multivariate graph limits (see Skeja and Olhede (2024)). These can be represented as a limiting object known as a decorated graphon (see Dufour and Olhede (2024)), where a more complex object is represented by a generalization of a graph limit. The choice of object will depend on the particular application at hand, and we will discuss the particular case of coma patients (Verdeyme et al (2025)). We will discuss choices of parameterization, and choices of summary statistics for particular models, highlighting both the benefits and challenges of multiplex observations.
  • Igor Prünster: Ties, Regularity, and Borrowing of Information in Dependent Nonparametric Priors Data from distinct, yet related, populations require a structured notion of probabilistic invariance, with partial exchangeability as the natural choice. Over the past two decades, many dependent nonparametric priors, including hierarchical, nested, and additive processes, have been proposed in this setting. Despite their widespread use in Statistics and Machine Learning, a unifying framework remains elusive, leaving key questions about their learning mechanisms unanswered. This talk focuses on multivariate species sampling processes, which generalize Jim Pitman’s species sampling construction to multiple related populations. They are characterized by a partially exchangeable partition probability function, encoding the induced multivariate clustering structure. Within this framework, borrowing of information across groups is governed exclusively by shared ties. This perspective also yields a taxonomy of dependent processes based on a notion of regularity and on the relationship between correlation, independence, and full exchangeability. This provides new insights into their learning mechanisms, including a principled explanation for the correlation structure observed in existing models, and indicates constructive routes towards richer forms of borrowing.
  • Angelina Roche: Minimax rates and adaptive estimation for prediction in linear models with functional data
    Functional data analysis is a relatively recent field of statistics specializing on data that can be modeled as random curves. In many situations, the aim is to predict a quantity (that can be e.g. scalar or even another functional data). It is now well known that leasts-squares estimators overfit the data in high-dimensional and infinite-dimensional contexts.
    Classical estimation procedures such as Ridge, Lasso, and projection estimators, rely on a smoothing parameter that has to be appropriately tuned to achieve an optimal compromise between bias and variance. Typically, the optimal value of this parameter depends on unknown quantities, such as the regularity of the estimated function. Selecting this parameter in a data-driven way is important for obtaining efficient, adaptive estimators. The aim of this talk is to present some recent work on this topic and highlight some open questions on adaptive estimation arising from the functional nature of the data.
  • Tenguao Wang: Coverage correlation: detecting singular dependencies between random variables
    We introduce the coverage correlation coefficient, a novel nonparametric measure of statistical association designed to quantifies the extent to which two random variables have a joint distribution concentrated on a singular subset with respect to the product of the marginals. Our correlation statistic consistently estimates an f-divergence between the joint distribution and the product of the marginals, which is 0 if and only if the variables are independent and 1 if and only if the copula is singular. Using Monge–Kantorovich ranks, the coverage correlation naturally extends to measure association between random vectors. It is distribution-free, admits an analytically tractable asymptotic null distribution, and can be computed efficiently, making it well-suited for detecting complex, potentially nonlinear associations in large-scale pairwise testing.
  • Johannes Wiesel: Dependence Measures via Adapted Optimal Transport: Stability and Rates of Convergence
    Recently studied dependence measures, such as Chatterjee’s rank correlation, that characterize both independence and perfect functional dependence, provide a powerful framework for detecting nonlinear dependencies. However, these measures cannot be weakly continuous, which limits the applicability of classical plug-in estimators based on empirical distributions. This obstruction is natural, as such measures are defined via conditional distributions and not through their joint law alone.
    In this paper, we introduce an optimal transport-based mode of convergence that captures weak convergence of conditional distributions and restores continuity for a broad class of dependence measures. We relate this mode of convergence to the adapted Wasserstein distance, the Knothe-Rosenblatt distance and the d1-metric on copulas. Building on this perspective, we propose a copula estimator based on the adapted empirical measure and compare it with the classical rank-based checkerboard estimator. For both estimators, we derive O(N1/3)O(N^{-1/3})-rates of convergence with respect to metrics that capture conditional weak continuity. As a consequence, we obtain the same rates for plug-in estimators of several classes of dependence measures, including rank-based and rearranged dependence measures.
    This talk is based on joint work with Jonathan Ansari.

Veranstaltungsort:

Ruhr-Universität Bochum
Im Beckmannshof
https://b-im-beckmannshof.de/