Dipankar Sarkar's picture
๐Ÿ—๏ธ Building on HF

Dipankar Sarkar PRO

dipankarsarkar

AI & ML interests

Building the AI-native stack. Agents as infrastructure, safety as architecture, performance as plumbing. I publish the receipts: papers, datasets, demos.

Recent Activity

liked a dataset about 2 hours ago
rmems/llm-eval-flakiness-trajectories
repliedto kanaria007's post about 2 hours ago
โœ… Article highlight: Benchmark Publication Without Governance Inflation (art-60-274, v0.1) TL;DR: This article argues that a benchmark result is not a governance maturity claim. A score may be real, reproducible, and worth publishingโ€”and still say nothing by itself about safety, deployability, assurance, institutional quality, or platform maturity. 274 treats benchmark publication as a discipline of comparability, disclosure, lifecycle limits, and anti-inflation. Read: https://huggingface.co/datasets/kanaria007/agi-structural-intelligence-protocols/blob/main/article/60-supplements/art-60-274-benchmark-publication-without-governance-inflation.md Why it matters: โ€ข prevents measured results from being inflated into safety or maturity claims โ€ข separates historical results from current comparability โ€ข makes scope, freshness, omissions, and unsupported readings visible โ€ข allows honest publication without requiring full platform assurance โ€ข treats narrower wording as trust discipline, not underselling Whatโ€™s inside: โ€ข the publication triad: comparability, disclosure, and anti-inflation โ€ข bounded publication outcomes such as PUBLISHABLE, PUBLISHABLE_WITH_LIMITS, NOT_COMPARABLE, and NOT_PUBLISHABLE โ€ข benchmark publication profiles โ€ข comparability disclosure notes โ€ข public non-claims registers โ€ข inflation checklists for result-to-maturity, comparison-to-assurance, historical-to-current, and wording inflation Key idea: Do not say: โ€œthis system scored well, therefore it is mature, safe, or ready to deploy.โ€ Say: โ€œthis result was observed under this benchmark and comparability frame, remains valid within these lifecycle and disclosure limits, and does not support these broader governance claims.โ€ Better benchmark publication is not a louder score. It is a result that is harder to overread.
View all activity

Organizations

Skelf Research's profile picture Neul Labs's profile picture Cognisoc's profile picture Incredlabs's profile picture