DIMP
DIMP (De-identify, Minimize, Pseudonymize) provides de-identification, minimization, and pseudonymization for FHIR data, protecting patient privacy while keeping the data useful for research.
What DIMP Does
- Removes or masks identifying information (names, addresses, etc.)
- Generates consistent pseudonyms for patient identifiers
- Preserves clinical data (diagnoses, procedures, lab values)
Configuration
Add DIMP to your aether.yaml:
services:
dimp:
url: "http://your-dimp-server:32861"
pipeline:
enabled_steps:
- torch # or local_import
- dimp # Pseudonymize after import
jobs_dir: "./jobs"Running Pseudonymization
aether pipeline start aether.yaml your-crtdl.jsonAether will:
- Extract data from TORCH (or import from files)
- Send it to DIMP for dimping
- Save the protected data in the jobs folder
Output
Results are saved in:
jobs/<job-id>/
├── state.json # Job state and step status
└── dimp/
└── dimped_<name>.ndjson.zst # Pseudonymized data (one file per input; .ndjson when compression is disabled)CRTDL Preprocessing
DIMP requires certain attributes (like Patient.identifier) to be present in the extracted FHIR data. If your CRTDL query doesn't include these attributes, DIMP pseudonymization will fail.
CRTDL preprocessing automatically enriches your CRTDL with the required attributes before sending it to TORCH:
services:
crtdl_preprocessing:
enabled: true
enrichments:
- group_reference: "https://www.medizininformatik-initiative.de/fhir/core/modul-person/StructureDefinition/PatientPseudonymisiert"
create_if_not_exists:
group_name: "PatientPseudonymisiert"
attributes_to_add:
- attribute_ref: "Patient.identifier"
must_have: false
- group_reference: "https://www.medizininformatik-initiative.de/fhir/core/modul-fall/StructureDefinition/KontaktGesundheitseinrichtung"
attributes_to_add:
- attribute_ref: "Encounter.identifier"
must_have: falseThe create_if_not_exists option creates the group in the CRTDL if it doesn't already exist. This is useful for groups like PatientPseudonymisiert that may not be part of the original research query but are needed by DIMP.
Enrichment rules can also be loaded from an external JSON file. See CRTDL Preprocessing in the configuration reference for details.
Linked Groups
An added attribute can link other attribute groups with linked_groups. Give the profile URL of each group. Aether changes each profile URL into the id of the attribute group that has this group_reference:
attributes_to_add:
- attribute_ref: "Patient.identifier"
must_have: false
linked_groups:
- "https://www.medizininformatik-initiative.de/fhir/core/modul-fall/StructureDefinition/KontaktGesundheitseinrichtung"Aether does this after it applies all the rules. Thus a rule can link a group that a subsequent rule creates, and the sequence of the rules does not change the result.
If a profile URL agrees with no attribute group, the pipeline stops with an error that gives the group, the attribute, and the URL. Aether writes no enriched-crtdl.json. Correct the URL, or add a rule that creates the group.
Anonymization Configuration
The $de-identify operation of the FHIR-Pseudonymizer accepts the anonymization configuration with each request. With this configuration, you can change the anonymization rules without a restart of the service.
This needs FHIR-Pseudonymizer v2.34.0 or later.
Give the path on the command line:
aether pipeline start aether.yaml crtdl.json \
--anonymization-config anonymization.yamlYou can also put the path in the configuration file. The flag overrides it:
services:
dimp:
url: "http://your-dimp-server:32861"
anonymization_config: "/path/to/anonymization.yaml"If you give no path, the service reads its anonymization configuration from its own file.
Aether reads the anonymization YAML at the start of the dimp step and sends it with each request. The job keeps the path, thus aether pipeline continue sends the same file without the flag. If the file is empty, or Aether cannot read it, the dimp step stops before it sends a request.
Earlier versions of Aether used the services.dimp.experimental_v3 block for this. That block is removed. A configuration file that still has it stops with an error. Move the path to services.dimp.anonymization_config.
Key Derivation
The cryptoHash method of the FHIR-Pseudonymizer needs a key. A different key for each project keeps the pseudonyms of the projects separate. But then you must generate and store one secret for each project.
The FHIR-Pseudonymizer can instead derive the key of each project from one master key. It uses HKDF (RFC 5869) with a derivation context. Key derivation needs FHIR-Pseudonymizer v2.30.0 or later.
Generate one master key for all projects:
openssl rand -hex 32Set this master key on the FHIR-Pseudonymizer server:
# compose.yaml of the FHIR-Pseudonymizer
environment:
Anonymization__CryptoHashKey: "<master key>"Then give each project a different context in the anonymization YAML:
fhirVersion: R4
fhirPathRules:
- path: Resource.id
method: cryptoHash
parameters:
keyDerivationContext: "project-a"The FHIR-Pseudonymizer derives the hash key from the master key and the context. The same master key with the same context always gives the same key. Two different contexts give two independent keys. Thus one master key is sufficient for many projects.
Keep the master key on the server. Do not set parameters.cryptoHashKey in the anonymization YAML. A key in parameters has precedence over Anonymization__CryptoHashKey, and the secret then moves with each request to the service.
You can also set a context on the server with Anonymization__KeyDerivationContext. A parameters.keyDerivationContext in the YAML has precedence over it. With the experimental v3 endpoint, Aether sends the YAML with each request. Thus the anonymization YAML of the project selects the key, and the server keeps only the master key.
The pseudonyms change if the master key or the context changes. Data that you send after a change does not agree with data from before the change. Use a stable context for each project.
The encrypt method uses the same mechanism for encryptKey. The FHIR-Pseudonymizer derives the hash key and the encryption key independently of each other.
Large Bundles
For large datasets, Aether automatically splits bundles before sending to DIMP:
services:
dimp:
url: "http://your-dimp-server:32861"
bundle_split_threshold_mb: 10 # Split bundles larger than 10MB
timeout: 30s # Timeout of one pseudonymization request