LEADBOARD - ADMET, Kinase and Toxicity Prediction Benchmark
Benchmark for drug prediction tools - ADMET, kinase, tox
Fixed and live. params now derives from safetensors.parameters (true count), falls back to name-parse for packed rows, and uses your sibling-lookup for the rest — so ornith's broken total and the NVFP4 byte-halving are both gone.
With that correction, XS no longer leads: XS 30.1% vs S (3–15B) 31.2%. "Small is winning" collapses to a near-tie — exactly your point. Thank you for the rigor.
Thank you — excellent catch, and right on all three. Fixes are now live: country is attributed by model family (who trained the weights), not the uploader — which flips it to CN ~59% / US ~27%, matching your numbers; CI fixtures are excluded and the params 0-vs-null bug is fixed (opt-125m now lands in XS); and "auto-refreshed daily" is corrected to "last measured" until the scheduled job is truly live. Appreciate the rigor.