ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures
Abstract
Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.
Community
π ALD/E-ImageMiner is an expert-annotated multimodal benchmark for understanding scientific figures from atomic layer deposition and atomic layer etching (ALD/E), covering both experimental and simulation studies. It contains 1,951 figures from 205 research papers.
Evaluate your models and submit results to the four CodaBench leaderboards:
π€ Access the complete ALD/E-ImageMiner benchmark dataset on Hugging Face
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- MSUE: Multi-Modal Soccer Understanding Expert (2026)
- SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation (2026)
- Benchmarking Deep Learning Approaches for AEC Engineering Drawing Layout Detection and Information Extraction (2026)
- Improving Reasoning in Vision-Language Models via Perception Verified Self-Training (2026)
- CIAN: Multi-Stage Framework for Event-Enriched Image Captioning via Retrieval-Augmented Generation (2026)
- Hierarchical Evidence-Driven Reasoning for Long Document Understanding (2026)
- ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.26848 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 2
SciKnowOrg/ALD-E-ImageMiner
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper