laya-browser β laya fine-tuned as a browser-agent decision head (drop-in replacement for TypeSafe Jev)
46-second demo (mp4): v10s drives headless Chromium through jev-ultrafast, one encoder pass per step; recorded with code/apps/make_demo.py.
laya (convaiinnovations/laya) is a non-autoregressive "System 1" decision model:
one bidirectional encoder pass answers several typed questions (choice / score / noul) with calibrated probabilities, no text generation.
Out of the box it is near chance at browser decisions ("which element should I click for this goal?" β top-1 0.10 among ~45 candidates).
This repo is what it took to turn it into a usable decision head for browser-use/jev-ultrafast,
whose /v1/systemone request format is identical to laya's predict(state, questions). Everything was done locally on one RTX 4070 Ti SUPER (16 GB),
no paid API: the text helper and the DAgger teacher are a local Qwen3-8B-AWQ served by sglang.
Progress update (2026-09-26): v12sβv14s, a held-out suite, and a data-leak correction
New checkpoint v14s/ (mmBERT-base 322M, format v3, run the server with a 60-option chunk threshold:
python apps/systemone_server.py 8791 /path/to/v14s 60).
| v10s | v12s | v13s | v14s | |
|---|---|---|---|---|
Suite B: 18 tasks on 18 sites that appear in no training source (Γ3, code/apps/browser_suite_b.py) |
β | 94 % (34/36, Γ2) | 93 % (50/54) | 100 % (54/54) |
| Suite A: the original 16 tasks (Γ3, patched harness) β contaminated, see below | 62 % | 75 % | 75 % | 81 % (39/48) |
| Held-out long-horizon synthetic tasks (flight / hotel / shop, 8 each) | β | β | β | 0 / 0 / 50 % |
Suite B checks the held-out claim at start-up against results/round3/train_domains.json (every domain of every training source,
682 domains). Its tasks are mostly 1β2 steps: navigation, site search, man pages, RFC / package search, a shop's search, page 2 and
sort-by-price. Suite A still fails books-page2 (0/3), Google Flights (0/3) and is flaky on two search tasks.
Data leak found and removed. dagger_cases.jsonl (the "177 on-policy DAgger corrections" in the table below, 331 cases by
now) came from running the Qwen teacher on the suite itself: 152 of its cases are suite-A goals, often with the teacher's wrong
label (e.g. "Go to page 2 of the catalogue" β click the first book), oversampled Γ5. Every checkpoint from v10 to v14s was trained on
it, so all suite-A numbers in this README, including v10s's 62 %, are contaminated (and books-page2 failing 0/3 is partly
the model faithfully reproducing a wrong label). Suite B is clean. From v15s on, build_items.py drops every DAgger case and every
case recorded on a suite A/B start page.
What changed in the harness (code/jev-ultrafast.patch, applies to jev-ultrafast 1231850). Half of the original suite
failures were harness bugs, not model errors:
- a
PRESS_ENTERcontrol, offered while a non-empty text field is focused (arXiv's search overlay has no submit button); - elements covered by an unrelated element (modals, overlays, banners) are no longer offered; a target covered by its own ancestor is clicked through;
- observation retries while the document is navigating; a choice that fails 3Γ is excluded, as is the same action repeated on the same URL (reload loops);
- an optional sub-goal planner (
JEV_PLANNER=1: the text model splits the task once, DONE advances to the next sub-goal, a text-model check rejects a sub-goal DONE the page does not show). Server side: wide choices are split into chunks above 60 options (a 160-link page gave each label ~4 tokens and a flat distribution), and a one-option choice no longer crashes the fast path.
What changed in the data (all collected in the real browser, no suite pages):
rollouts2.py: scripted search βPRESS_ENTER/submit β DONE, "open the article about X", scroll-to-next-page, scroll-then-click,<select>with intent-style goals ("Sort by price: low to high."), type-then-pick-a-suggestion; 2.3k cases on 206 sites.convert_nnetnav.py: 8.9k steps from stanfordnlp/nnetnav-live (element tables capped at 45, held out by site).- target lists sampled to the server's chunk width (β€ 40) and final-round width (2β6) so labels are never truncated in training.
Where it breaks: long horizons. On held-out synthetic forms (below) v14s ends flight and hotel searches after 1β2 steps (DONE too early) or loops. Jev's own Google Flights run needs 11 actions: one-way, type origin β pick the suggestion, type destination β pick the suggestion, open the calendar, pick the day, Done, Search.
In progress: v15s (clean retrain from the base, no DAgger, no suite pages) adds:
code/finetune/webgym/: a local synthetic web environment (flight / hotel search forms, shop listings) with randomized widget implementations β autocomplete comboboxes inline or behind a trigger button, month-grid date pickers with Done/Apply, radio / segmented / native / custom dropdowns, steppers and popups, cookie banners, forms pushed below the fold. A scripted expert reads the page's own state, acts through jev's observations (so the format is exactly what the agent sees), and records each step twice: under the full task and under the current sub-goal, plus a verified sub-goal DONE whenever a sub-goal is satisfied. 12 % harmless detours teach recovery (including going back after a wrong submission). 1,408 successful episodes, 32k cases; seeds900000 + 10Β·iare held out foreval_gym.py.annotate_subgoals.py: Mind2Web trajectories segmented into planner-style sub-goals by the local Qwen (trajectory-conditioned, Plan-and-Act style), 5.1k sub-goal-conditioned steps. Sub-goal DONE labels come only from webgym, where they are checked against page state. Results will be added here when the run finishes.
What changed relative to the original laya
| original laya (typed-decisions) | this repo | |
|---|---|---|
| browser decision quality (16 real tasks Γ 3 runs, corrected checks) | 0 % | v10s: 62 %, v11s: 56 %, v10: 50 % (10 tasks pass 3/3, 6 fail 3/3; see below) |
| element top-1 on held-out pages (2,734 decisions, ~45 candidates) | 0.10 | 0.66 (v10), 0.63 (v10s), 0.62 (v11s) |
| operation accuracy (CLICK / TYPE_TEXT / SELECT / DONE) | 0.54 | 0.88β0.89 |
| latency per browser step (3 questions, 30β65 candidates) | 50β200 ms | 41β50 ms (v10), 17β23 ms (v10s / v11s) |
| backbone | ModernBERT-large 421M | v10: same; v10s / v11s: mmBERT-base 322M |
| input format | jev's state verbatim (element table as JSON inside the state, truncated by the 1024-token window) | format v2/v3: elements live only in the option list (full label + role + current value), state keeps title / URL / history / 1.2β1.5k chars of text, head_max_len 512 β 768 |
| training data | LocalLLaMA/typed-decisions | 5,244 reverse-generated goals on 421 crawled pages (Qwen writes "the goal a user would state to need this element"), 700 real DONE states (clicks actually executed), 659 step-2 negatives, Mind2Web train (7,296 steps, candidates re-rendered as an element table), 177 on-policy DAgger corrections (later found to contain suite-A tasks, see the progress update) |
| training | β | laya's RLCD recipe (noisy-logit policy gradient + soft CE), single GPU, no gradient checkpointing, 4 epochs (~2 h for v10, ~1 h for v10s), post-hoc temperature |
| inference | HF eager + autocast | optional TileLang fast path (PR #25 to laya): fused GEMM/GEGLU/LayerNorm/RoPE, sliding-window flash attention, bf16-resident weights, CUDA graphs β 4β5Γ lower per-call latency, identical answers |
Things that did not work (so you don't repeat them)
- Templated DONE goals ("Open the page titled X, stop once it is open") leak phrasing: the model learns stop when β DONE. DONE samples must be real landing pages after an executed action.
- If every DONE sample has exactly one prior action and every click sample has none, the model learns any history β DONE. Add mid-task negatives (step-2 goals on landing pages).
- Mind2Web alone kills DONE / TYPE_TEXT (no DONE there, CLICK dominates): re-weight rare operations (DONE Γ4, TYPE_TEXT/SELECT Γ3).
- Cutting page text to 3,000 chars saved nothing (the sequence is dominated by the head) and cost 0.04 top-1.
torch.compileon variable-length batches recompiles per shape: 6Γ slower.- Confidence-gated escalation to Qwen3-8B (System 2) made things worse (58 % β 42 %): on these pages the fine-tuned 322M/421M model is a better decider than an 8B general LLM. Use a stronger System 2 or none.
- jev's DOM reader hides password fields by design (login tasks are impossible) and never sees collapsed menus (Wikipedia's "Random article").
What still fails
The suite is bimodal: 10 tasks pass 3/3 (category / tab / page navigation, checkbox, <select>, HN pages, DuckDuckGo search in some runs) and 6 fail 3/3:
"type then submit / pick a suggestion" flows (Wikipedia search Γ2, arXiv), pagination that needs a scroll first (the model clicks the first
visible item instead), and Google Flights. v11s added 682 scripted scroll / search-submit / select trajectories (code/finetune/rollouts.py):
SELECT and DuckDuckGo improved, the scroll case did not β run-to-run variance on live sites (HN front page changes, DDG 50x pages) is larger
than the v10s β v11s difference, so treat the two as equivalent.
Teachers vs the fine-tuned student (80 held-out decisions, same element tables)
| decider | op acc | target top-1 | s / decision |
|---|---|---|---|
| Ternary-Bonsai-2-27B (local, thinking off) | 0.825 | 0.554 | 1.56 |
| Bonsai-27B, thinking budget 300 tokens | 0.861 | 0.603 | 4.7 |
| Bonsai-27B, thinking unrestricted | 0.625 | 0.474 | 11 |
| Qwen3-8B-AWQ (35 % of requests failed, survivors only) | 0.904 | 0.617 | 0.38 |
| laya v11s (322M, this repo; full 2,734-case set) | 0.890 | 0.623 | 0.021 |
A 27B general model with a short thinking budget matches the 322M fine-tuned student at 200Γ the latency; neither 8B nor 27B is a useful DAgger teacher or System-2 fallback here. Further gains need a stronger teacher or more targeted trajectories.
Files
v14s/ mmBERT-base 322M, format v3, head_max_len 768 (best: suite B 100 %; serve with chunk threshold 60)
v10/ ModernBERT-large 421M, format v2, head_max_len 768 (best held-out top-1)
v10s/ mmBERT-base 322M, format v3, head_max_len 768 (17β23 ms per step; best on the live suite)
v11s/ v10s data + 682 scripted scroll / search-submit / select trajectories (equivalent to v10s within noise)
code/ finetune pipeline, laya systemone server, task suite, TileLang kernels, jev-ultrafast patch
results/ per-run suite JSONs and logs behind every number above (round3/: v12sβv14s, suite B, webgym baseline)
Each checkpoint is a laya checkpoint directory (model.safetensors, encoder/, tokenizer/, rl_agent_config.json); the config records
laya_fmt and head_max_len_train so the server applies the matching input format automatically.
Use
huggingface-cli download cklxx/laya-browser --local-dir laya-browser
cd laya-browser/code && uv sync --extra fast # pinned uv.lock (Python 3.12, torch 2.11, tilelang 0.1.14)
uv run python verify.py v10s # downloads v10s if needed, answers one recorded browser step
uv run python verify.py v10s --fast # same through the TileLang fast path
Verified from a clean environment on 2026-09-21 (RTX 4070 Ti SUPER): TYPE_TEXT β [2] Search Wikipedia (searchbox), 35 ms per step
stock / 28 ms with the fast path on a 65-option, 2.5k-token step. Extras: --extra data (Mind2Web conversion, dataset eval),
--extra browser (live suite / crawling; also needs jev-ultrafast with code/jev-ultrafast.patch applied and a Chromium with
--remote-debugging-port=9222).
import laya
agent = laya.load("laya-browser/v10s") # a local laya checkpoint dir
agent.cfg["head_max_len"] = agent.cfg["head_max_len_train"]
# state / questions exactly as jev-ultrafast's model.choose() builds them, after the format-v3 transform in code/apps/systemone_server.py
result = agent.predict(state, questions)
As a TypeSafe replacement for jev-ultrafast:
# in code/: laya systemone-compatible server (format transform + optional gating + DAgger logging)
python apps/systemone_server.py 8791 /path/to/laya-browser/v10s 999
# in jev-ultrafast (apply code/jev-ultrafast.patch): TYPESAFE_BASE_URL=http://127.0.0.1:8791
code/apps/browser_suite.py runs the 16-task real-browser suite with automatic outcome checks (REPEATS=3).
Reproduce
code/finetune/README.md documents every step (crawl β reverse-generate goals β execute clicks for DONE β step-2 β Mind2Web conversion β
DAgger β build β train β calibrate β eval β suite) with the exact scripts (run_v10.sh, run_v10s.sh, run_final.sh) and all intermediate
numbers from v1 to v10s.
GPU cost
v10s: ~0.65 GB weights, ~1.5 GB VRAM resident with CUDA graphs, 17β23 ms per 3-question browser step, 3 ms for a single-question call. v10: ~0.85 GB weights, ~1.8 GB VRAM, 41β50 ms per step. The Qwen text helper (only needed for TYPE_TEXT values) is separate.
License
Apache-2.0, same as laya. Mind2Web is used under its own license for training only.
