Skip to content

Architecture ​

FLAME consists of two parts: one central Hub and one Node at every participating site. The Hub coordinates, the Nodes compute. Patient data stays at the site, only analysis code travels to the data and only aggregated results travel back.

The graphic gives a rough overview of the execution flow of an analysis. The individual services are listed on the Components page.

Hub and Node ​

HubNode
Runsonce, centrallyonce per site (hospital, university, ...)
Purposecoordinationexecution on local data
Holdsusers, projects, analysis code, analysis images, resultsthe site's data stores and running analyses
Sees patient datanoyes, inside the site only
Used byresearchers and administratorsnode administrators and data stewards

Hub ​

The Hub is the single place researchers work with. Its services:

  • User Interface to create projects and analyses, review them and download results.
  • Core as the main backend, managing all resources and triggering commands and events.
  • Authentication and authorization (Authup), organized in one realm per organization.
  • Analysis building, which combines the approved code with a master image into a container image.
  • Container image registry (Harbor) with a separate, private project for every node.
  • Message Broker, which relays the messages nodes send to each other.
  • Storage for code, intermediate results and final results.

Node ​

A Node is installed on a Kubernetes cluster inside the network of a site. Its services:

  • Node UI to configure data stores and to start, stop and monitor analyses.
  • Hub API Adapter, the backend of the Node UI and the connection to the Hub.
  • Pod Orchestration, which starts every analysis in its own isolated pod.
  • Data access through an API gateway (Kong), which exposes exactly the data stores of a project to its analyses.
  • Message Broker and Result Service, the node side of the communication with other nodes and the Hub.

Lifecycle of an analysis ​

  1. Project: a researcher creates a project in the Hub and selects the nodes that should take part. Every selected node approves or rejects the project.
  2. Data stores: at every approving site, an administrator or data steward connects the requested data (a FHIR server or an S3 bucket) to the project as a data store.
  3. Analysis: the researcher creates an analysis within the project, uploads the code and selects a master image with the required dependencies.
  4. Review: the administrator of every node can download and read the code, and approves or rejects the analysis.
  5. Build and distribution: after the approvals, the Hub builds one container image from master image and code and pushes it into the registry project of each node.
  6. Execution: every node pulls its image. The node administrator starts the analysis in the Node UI. The analysis runs in an isolated pod and can only reach the data stores of its own project.
  7. Aggregation: the analyzing nodes send their results to the node with the aggregator role, which combines them. For iterative methods such as federated learning this repeats over several rounds.
  8. Results: the aggregator submits the final result to the Hub storage, where the researcher downloads it.

The same steps from the perspective of each role are described in the User and Admin guides.

Communication ​

  • Nodes connect to the Hub, never the other way round. All connections are initiated by the Node as outbound HTTPS. A Node does not have to be reachable from the internet.
  • Nodes do not talk to each other directly. Messages and intermediate results are relayed through the Hub.
  • What is relayed is encrypted end to end. Messages and intermediate results are encrypted for the receiving nodes with the node key pairs, so the Hub cannot read them.
  • Analyses are isolated. An analysis container has no internet access and no access to other services of the Node. It only reaches its project's data, the message broker and the result service through a dedicated proxy.

Details are documented in Security & Operations.

Where data lives ​

DataLocationLeaves the site
Patient data (FHIR resources, files)data stores of the siteno
Analysis codeHub storage, and inside the analysis imageyes, it is distributed to the nodes
Analysis imageHub registry, pulled by the nodesyes, it is distributed to the nodes
Messages and intermediate resultsrelayed and stored by the Hub, end-to-end encryptedyes, readable only by the receiving nodes
Local intermediate results and checkpointsstorage of the nodeno
Final resultsHub storageyes, downloadable by the researcher

Because final results leave the federation, an analysis must only return aggregated, non-identifying values. This is what the code review by the node administrators is for.

Deployment ​

Hub and Node are both deployed with Helm charts on Kubernetes. See the Deployment guide.