ChengyiX/KLIK-Bench
Viewer • Updated • 58 • 209 • 1
Public datasets for memory-grounded and command-line AI-agent evaluation. Research artifacts only; not commercial Klik releases or product validation.
Note Persona-aware multi-tool orchestration evaluation covering memory, preferences, recovery, and confirmation-boundary tasks.
Note Command-line tool-orchestration evaluation with deterministic mock backends and outcome, efficiency, recovery, and consistency metrics.