EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
Abstract
Large language models can generate interactive web interfaces, but reliable generative UI requires maintaining an executable artifact as user requests evolve. We introduce EvoGenUI-Bench, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information presentation, executable interaction, and tool-grounded external state. We execute generated artifacts in a browser and evaluate them using screenshots, source and DOM evidence, actor traces, and runtime logs. Beyond turn-level and episode-level success, we measure cross-turn retention with Adjacent Pass Retention. Across eight models, even the strongest achieves 74.9% Turn Pass while completing only 37.3% of five-turn episodes; APR further falls to 52.4% on tool-grounded tasks. Diagnostic analysis shows that presentation failures center on information architecture, interaction failures on derived-state propagation and affordance binding, and tool-grounded failures additionally involve external-state grounding and requirement decomposition. These results reframe generative UI evaluation from judging isolated outputs to testing whether interface behavior, derived state, external state, and assistant claims remain synchronized as the artifact evolves.
Community
Can generative UI assistants keep one executable interface correct as user requirements evolve across multiple turns?
EvoGenUI-Bench introduces 150 five-turn tasks (750 turns) spanning information presentation, executable interaction, and tool-grounded external state. Generated artifacts are executed in a browser and evaluated using screenshots, source/DOM evidence, actor traces, and runtime logs.
Across eight models, the strongest reaches 74.9% Turn Pass but completes only 37.3% of full five-turn episodes. We also introduce Adjacent Pass Retention (APR) to measure whether behavior that worked at one turn survives the next update. The results show that cross-turn maintenance—not just one-shot visual quality—is a central bottleneck for reliable generative UI.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows (2026)
- MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents (2026)
- StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows (2026)
- DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation (2026)
- BTS-AgentBench: A Deterministic, Replayable Pipeline from Read-Only Telemetry Logs to Agent Benchmarks (2026)
- UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation (2026)
- Repo2Skill-Evo: Repository Skills Go Stale in Silence (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.29387 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper