AI Safety Evaluation Datasets Collection Datasets and viewer resources for studying evaluation awareness, instrumental behaviour, sabotage, and monitoring in AI agents. • 3 items • Updated 3 days ago
Agent Trajectories Collection Agent trajectories from the AI Safety and Alignment Group’s benchmarks, released for analysing agent behaviour, capabilities, and safety. • 4 items • Updated 3 days ago
Pythia Scaling Suite Collection Pythia is the first LLM suite designed specifically to enable scientific research on LLMs. To learn more see https://github.com/EleutherAI/pythia • 18 items • Updated Feb 26, 2025 • 35
ICML 2026 Reproduction Collection Public evidence-led ICML 2026 reproduction logbooks, bundles, artefacts, and primary paper sources. • 55 items • Updated 4 days ago
Running Reproduction: SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs 🎯 Explore simulation logs, traces, and workspace files
Running Reproduction: SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs 🎯 Explore simulation logs, traces, and workspace files
ICML 2026 Reproduction Collection Public evidence-led ICML 2026 reproduction logbooks, bundles, artefacts, and primary paper sources. • 55 items • Updated 4 days ago
Running Reproduction: CauchyNet: Compact and Data-Efficient Learning Using Holomorphic Activation Functions 🎯 Explore project code, traces, and workspace in an interactive logbook
Running Reproduction: CauchyNet: Compact and Data-Efficient Learning Using Holomorphic Activation Functions 🎯 Explore project code, traces, and workspace in an interactive logbook
ICML 2026 Reproduction Collection Public evidence-led ICML 2026 reproduction logbooks, bundles, artefacts, and primary paper sources. • 55 items • Updated 4 days ago
Running Reproduction: A hitchhiker's guide to Poisson gradient estimation 🎯 Browse code, traces, and workspace in an interactive logbook