Model Card for ESFM/ESFM_s_wm_ri
Final released random-initialization ablation checkpoint with masked ERA5 training. It isolates the effect of starting ESFM without Aurora distillation or CMIP6 pretraining.
Checkpoint selection: Use for initialization-ablation studies and reproducibility of the random-start experiment.
Model Details
- Developed by: The ESFM research team, with the full contributor and author lists linked below.
- Shared by: ESFM on Hugging Face
- Model type: Final random-initialization comparison checkpoint (ESFM_s,ri); modified 3D Swin-UNet encoder-decoder
- Model size: Approximately 115 million parameters
- Masking protocol: Variable, pressure-level, and spatial masking
- Forecast lead time: 6 hours
- License: MIT
- Repository: https://huggingface.co/ESFM/ESFM_s_wm_ri
Model Sources
- Code: https://github.com/swiss-ai/ESFM
- Paper: https://arxiv.org/abs/2605.00850
- Project page: https://swiss-ai.github.io/ESFM/
The paper is currently available as an arXiv preprint.
Uses
Direct Use
Use for initialization-ablation studies and reproducibility of the random-start experiment.
Downstream Use
Research on pretraining and initialization effects for heterogeneous Earth-system models.
Out-of-Scope Use
Not the recommended general ESFM checkpoint; use ESFM_s_wm for the default knowledge-distilled model.
Bias, Risks, and Limitations
The manuscript reports lower accuracy than the knowledge-distilled initialization. The model also inherits ERA5 biases and is not designed as an operational forecasting system.
All ESFM checkpoints are research artifacts. Users should validate forecasts for their variables, regions, seasons, lead times, missingness pattern, and decision context. Do not use the model as the sole basis for safety-critical decisions.
How to Get Started
The checkpoint is not packaged as a Hugging Face Transformers from_pretrained model. Construct the ESFM architecture with the matching repository config, then load the state dictionary. The released notebook contains the complete download, model-construction, normalization, and inference workflow.
git clone https://github.com/swiss-ai/ESFM.git
cd ESFM
# Open notebooks/inference_ESFMs_on_ERA5.ipynb
In the notebook, set:
EXPERIMENT_NAME = "ESFM_s_wm_ri"
To download the weights directly:
from huggingface_hub import hf_hub_download
model_name = "ESFM_s_wm_ri"
weights_path = hf_hub_download(
repo_id=f"ESFM/{model_name}",
filename=f"{model_name}.safetensors",
)
print(weights_path)
Set EXPERIMENT_NAME = "ESFM_s_wm_ri" in the released inference notebook, or run the matching configs/config_ESFM_s_wm_ri.yaml.
Training Details
Training Data
WeatherBench2 ERA5 at 0.25-degree resolution, trained on 1979 through 2020 and evaluated in the manuscript on held-out ERA5 forecasts.
Dataset preprocessing and the exact variable registry are documented in the ESFM repository and preprint.
Training Procedure
Continues training from ESFM_s_wm_ri_pre for 10,000 steps. This experiment was run on 16 GPUs.
- Training objective: Six-hour forecast learning, as specified above
- Nominal architecture: ESFM small, approximately 115M parameters
- Software environment: PyTorch/Lightning in the released NVIDIA PhysicsNeMo 25.03 container; lightning==2.5.1 is pinned in the Dockerfile
- Training regime: Lightning
precision="32-true"with FP32 parameters and optimizer state; selected model forward operations use CUDA BF16 autocasting throughtorch.autocast(dtype=torch.bfloat16).
Evaluation
The manuscript compares random, CMIP6, and knowledge-distilled initialization under the same masked ERA5 forecasting setup and reports the random-start model as the weakest of the three. See the preprint for current details.
The manuscript uses held-out temporal data and reports task-appropriate metrics: latitude-weighted MAE and Pearson correlation for gridded deterministic forecasts, relative MAE for MODIS comparisons, station metrics for station models, and CRPS for ensembles. Detailed values are intentionally not copied into this card.
Technical Specifications
ESFM retains Aurora's 3D Swin-UNet backbone and adds variable-specific tokenization, axial attention across variables, perceiver aggregation across variables and pressure levels, learnable NaN tokens for missing patches, resolution-specific tokenizers where configured, and a decoder queried at target pressure levels. The small configuration uses a 256-dimensional embedding and approximately 115M parameters.
Environmental Impact
- Hardware type: NVIDIA GH200 systems with four GPUs per node. This experiment was run on four nodes, totaling 16 GPUs.
- Total training time: 24 hours
- Compute location: Training used CSCS Alps infrastructure.
Citation
@misc{ozdemir2026esfm,
title={Earth System Foundation Model (ESFM): A unified framework for heterogeneous data integration and forecasting},
author={Firat Ozdemir and Yun Cheng and Salman Mohebi and Fanny Lehmann and Simon Adamov and Zhenyi Zhang and Leonardo Trentini and Dana Grund and Oliver Fuhrer and Torsten Hoefler and Siddhartha Mishra and Sebastian Schemm and Benedikt Soja and Mathieu Salzmann},
year={2026},
eprint={2605.00850},
archivePrefix={arXiv},
primaryClass={physics.ao-ph},
url={https://arxiv.org/abs/2605.00850}
}
More Information
Model Card Contact
Firat Ozdemir: firat.ozdemir@sdsc.ethz.ch
Model tree for ESFM/ESFM_s_wm_ri
Base model
ESFM/ESFM_s_wm_ri_pre