Explanatory Data Analysis group
Research
We develop algorithms and theory that help people discover what matters in complex data, understand cause and effect, and make trustworthy predictions and decisions.
Our fundamental research has two complementary themes. Discovering what matters is about finding compact, informative, and understandable structure in data. From causes to decisions is about learning how systems work, reasoning about interventions, and choosing effective actions. Together, these themes form our approach to explanatory data analysis.
Discovering what matters
Large and complex datasets can contain many apparent patterns, but only a small fraction are informative, reliable, and useful. We develop methods for exploratory data mining and interpretable machine learning that identify this relevant structure and express it in forms people can understand. These include probabilistic rule sets, decision trees, histograms, patterns, subgroups, and explanations of anomalous observations.
Where possible, we favour models that are interpretable by construction: their structure, predictions, and uncertainty can be examined directly, rather than explained only after the fact. Although much of our work starts from tabular and other structured data, we increasingly extend these ideas to event sequences, dynamic graphs, time series, and spatiotemporal data.
The minimum description length (MDL) principle is central to much of this work. Based on this principle we view learning as finding the model that compresses the data best, providing a principled way to balance goodness of fit against model complexity. This allows us to learn concise probabilistic models and explanations without relying on arbitrary regularisation choices. The same data may support several models that perform almost equally well. We investigate this set of alternatives—the Rashomon set—and how background knowledge, expert feedback, and interactive exploration can help identify models that are not only accurate, but also meaningful and useful.
Explainable anomaly detection is part of this broader programme: detecting that an observation or process behaves unusually is often only the beginning. We want to identify what makes it unusual and present that evidence in a form that supports investigation and action. Where useful, we combine interpretable modelling principles with modern scalable learning techniques, while keeping the explanation and its reliability central.
Selected recent work
Probabilistic Truly Unordered Rule Sets. Journal of Machine Learning Research, JMLR, In press. |
|
Scalable, Explainable and Provably Robust Anomaly Detection with One-Step Flow Matching. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS 2025), 2025. |
|
Conditional Density Estimation with Histogram Trees. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS 2024), 2024. |
From causes to decisions
Predictions and associations alone cannot tell us what will happen when we intervene. We therefore also develop causal learning methods that aim to uncover how variables influence one another, estimate the effects of possible actions, and distinguish genuine causal relationships from correlations caused by hidden or shared factors.
Real data rarely comes from one clean and fully controlled experiment. It may combine observations from different environments, populations, or experimental conditions; some intervention targets may be unknown, and important factors may be unobserved. We study when causal effects and structures can nevertheless be identified, and develop algorithms that exploit heterogeneity, higher-order statistical information, and structural causal models to do so.
We also study how causal knowledge can guide decisions. This includes choosing informative experiments, targeting interventions, and learning sequences of actions through reinforcement learning and optimisation. The aim is to move from describing and predicting a system to understanding which actions are likely to achieve a desired outcome, while representing uncertainty and the assumptions on which those conclusions depend.
Selected recent work
Causal Effect Identification in Heterogeneous Environments from Higher-Order Moments. In: Proceedings of the Conference on Uncertainty in Artificial Intelligence (UAI 2025), 2025. |
|
MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters. In: Proceedings of the International Conference on Machine Learning (ICML 2025), 2025. |
|
Hierarchical Reinforcement Learning with Targeted Causal Interventions. In: Proceedings of the International Conference on Machine Learning (ICML 2025), 2025. |
|
Learning Unknown Intervention Targets in Structural Causal Models from Heterogeneous Data. In: International Conference on Artificial Intelligence and Statistics, pp 3187-3195, PMLR, 2024. |
Where the themes meet
Learning Subgroups with Maximum Treatment Effects without Causal Heuristics, presented at AAAI 2026, is a direct example of how our two themes reinforce one another. It combines interpretable subgroup discovery with causal-effect estimation to find understandable groups of subjects for whom a treatment or intervention has the greatest effect.
These results demonstrate how questions about interventions lead to a new pattern-discovery problem, while interpretable learning makes the resulting causal conclusions easier to inspect and use. It illustrates our broader goal of developing methods that reveal meaningful structure and turn it into evidence for action.
Learning Subgroups with Maximum Treatment Effects without Causal Heuristics. In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2026), 2026. |
Many of these fundamental questions originate in our work with researchers and organisations in other fields. See our interdisciplinary research for the challenges that inspire this work and the settings in which we develop it.