Real-Time Inference

Efficient and fast computation of non-equilibrium models via Directed Acyclic Graphs

Overview

Non-equilibrium inspired models, such as diffusion/flow-matching, are costly in terms of their inference time. Requiring significant acceleration to be even remotely viable in high energy particle physics data analyses. As such, NEEDLE leverages a combination of computational concurrency based methods paired with efficient mathematical methods for solving probability paths core to the continuous time non-equilibrium models - e.g. diffusion/flow-matching - in order to reduce the computational overhead and make diffusion/flow-matching models practically available.

Graph-based Inference

Pseudo-Surrogate

The NEEDLE CLI training DAG workflow enables users to train an ensemble of neural surrogates, representing k-folds, randomly seeded ensembles, systematics, and the unique estimators that comprise the statistical model of the data. This introduces the problem of how to combine the outputs of the neural surrogates into a single tensor that can be used as a single cohesive entity in any statistical analysis - i.e. a single likelihood value for each piece of data. To address this problem we build the idea of a Pseudo-Surrogate:

  • Mixture of Neural Surrogates: Dynamically combine by DAG inversion many neural-surrogates into a single neural network, or pseudo-surrogate, that can be handled via the needle-api as a single PyTorch model definition.
  • Aggregation Methods: The edges of the DAG can be configured at runtime via the needle-api to combine the sub-surrogates of the pseudo-surrogate as per the user request. For example, the k-folds can be averaged, ensembles can be combined based on 'best-of'all-models', and estimators can be a weighted sum.
An illustration of this sub-surrogate combination by DAG inversion is given below, for three modes of inferencing:

Accelerating Inference

Inference Modes

The DAG based inferencing model allows for a single pseudo-surrogate to be defined in a programatic fashion, that users can interact with via a single API call. Unfortunately, the pseudo-surrogate model will scale linearly with the number of neural surrogates in the DAG, but also non-linearly with the size of each individual network. Consequently, to make the NEEDLE toolkit practically useful in every day data analyses, the NEEDLE API for graph inferencing offers three modes of acceleration:

    • Sequential: Treat each neural sub-surrogate in the DAG as a separate PyTorch model that executes in series a 'forward pass' on an accelerator device orchestrated by the inverted DAG.
    • Parallel: Treat each neural sub-surrogate in the DAG as a separate PyTorch model that executes as a CUDA stream across the Stream Monitors of NVidia GPUs. Multiple streams can run concurrently, and across multi-gpus at the same time.
    • Vectorised: Batch all matrix operations across the neural sub-surrogates as single operations on the CUDA streams, allowing for concurrent layer wise execution all sub-surrogates simultaneously.

Performance

Preliminary benchmarks of inference runtime for each mode of operation. Select a network architecture to compare:

Runtime vs. network parameter count:

Runtime vs. number of networks in the DAG:

The benchmarks demonstrate the key advantage of concurrent acceleration for the pseudo-surrogate. The sequential mode of operation, whilst simple, is the slowest methodology when running on a single NVidia Tesla V100 card, due to the linear complexity scaling in either the number of networks in the DAG, or non-linear scaling in network parameters. The vectorized mode of operation offers substantial improvements in acceleration that substantially curbs the linear and non-linear scaling of the complexity problem as a function of network size (num. params) or the number of networks in the DAG, whilst using a single GPU. However, the parallel mode of operation offers scalability in terms of horizontal acceleration (across multiple GPUs), where in the above test 2 NVidia Tesla V100 GPUs were used to concurrently process the networks using individual CUDA streams. Whilst the most performant mode of operation this comes at the expense of more computational resources.

Back to Project Overview