Project Overview
Mission
The NEEDLE (NEural-basEd Diffusion Likelihood Estimations) project is a toolkit for Neural Simulation Based Inference methods. Its purpose is to be the training and inference engine that accelerates the deployment of existing and upcoming NBSI tools. It takes care of model orchestration, config management, and submission to batch systems for the High Energy Physics community.
Our work focuses on three core pillars: developing new neural-based surrogates, developing new likelihood/posterior inference methodologies within the diffusion/flow-matching domain, and creating an efficient data ingestion and orchestration system for highly parallelisable HTC/HPC systems. The work herein is the collaborative effort of particle physics groups within the ALICE, ATLAS, and CMS Collaborations at the European Center for Nuclear Research (CERN).
Software Stack
Core Technologies
Development Tools
- Version Control: GitHub + GitLab CI/CD pipelines
- Containerization: Docker, Singularity for HPC environments
- Testing: pytest, continuous integration workflows
- Documentation: mkdocs-materials & docstrings
Architecture Design
DAG-Based Orchestration
NSBI tools are seldom a single neural network that infers the likelihood (or adjacent quantities) in a single shot. In practice, the surrogate is rather composed of different sub-models that together predict the final quantity. In some cases, that number can reach into the hundreds of models, especially if techniques like ensembling are used to mitigate the biases coming from stochastic sources in the training.
NEEDLE uses a graph-structure to encode the relations between networks, a Directed Acyclic Graph (DAG). This DAG is also the natural representation of computational workflows in larger HEP analyses, with each step being an individual Task. Each model can be thought of as a single Task with clear inputs (training data) and outputs (trained model checkpoint). The NEEDLE framework uses a luigi-based workflow to orchestrate different Tasks in parallel.
Key Features
- Checkpointing: Automatic recovery from failures, unfinished Tasks are resubmitted
- Versioning: Full ML training tracking and experimentation for reproducibility with the MLflow library
Integration Points
- CERN computing grid (WLCG)
- DESY National Analysis Facility. Tested on NAF (htcondor) itself and MAXWELL (slurm)
- KIT
Machine Learning Research
Current Research Areas
Optimizing transformer model architectures to efficiently ingest irregular data sizes, in order to reduce memory footprint of ML training and inference tasks.
Recent Progress
Q1 2026 Milestones
- Release of the NEEDLE workflow managing framework
needle-sbi, which is available on GitHub - Demonstration of the NEEDLE framework on the FAIR Universe HiggsML Challenge dataset (using code by I. Elsharkawy, found here)
Upcoming Goals
- Presentation at CHEP 2026 Conference
- White paper with technical breakdown of the software