Covering Human Action Space for Computer Use: Data Synthesis and Benchmark Paper • 2605.12501 • Published May 12 • 17
From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering Paper • 2603.20193 • Published Mar 20 • 1
LLMSurgeon: Diagnosing Data Mixture of Large Language Models Paper • 2605.30348 • Published May 28 • 1
Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense Paper • 2602.09012 • Published Feb 9
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems Paper • 2604.14228 • Published Apr 14 • 25
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Paper • 2608.13560 • Published Aug 13 • 64
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting Paper • 2602.17645 • Published Feb 19