Instructions to use Ameame1002/RLCurator with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Ameame1002/RLCurator with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
- CuratorRL - replay-validated trajectory-repair curators for AppWorld agents
- Two curator architectures
- What is in this repository
- How these numbers were measured
- Dev results - equal-weight macro over three actors
- Dev results - per actor
- Each GRPO arm against its own initialization
- What we found
- Limitations - please read before citing a number
- How to load and run
- Provenance and integrity
- Citation
- License
- 中文摘要
- Two curator architectures
CuratorRL - replay-validated trajectory-repair curators for AppWorld agents
Twelve LoRA adapters for Qwen/Qwen3-8B that implement a curator: a model that reads a failed AppWorld tool-using trajectory and proposes a minimal repair.
failed ReAct trajectory -> { "step_id": <checkpoint to resume from>,
"repair_hint": "<= 2 sentences" }
The point of the project is how that output is scored. A repair is not accepted because a teacher model wrote it; it is accepted because a replay verifies it helped:
- restore the AppWorld environment snapshot saved at
step_idand re-execute the trajectory prefix so the Python interpreter namespace is rebuilt too; - HINT branch - append
Guidance: <repair_hint>and let the actor continue; - BASE branch - restore the same checkpoint and the same seed, append nothing, and let the actor continue (the actor is not told a hint is absent);
- judge both branches with AppWorld's ground-truth verifier over three replay seeds.
The matched base branch is what separates the hint caused the success from that checkpoint was recoverable anyway. The same replay signal is the reward used by the GRPO adapters in this repository.
Source repository: https://github.com/22856226/RlCurator
Two curator architectures
| architecture | modules | how the repair is produced |
|---|---|---|
| Factorized | selector + decoder (two adapters) |
the selector scores every legal checkpoint step and takes the argmax (earliest-step tie-break); the decoder then writes the hint conditioned on that step |
| Joint | one joint adapter |
a single free-form greedy generation emits the step_id and the hint together |
Both are trained in two steps. Step 2 (supervision) fits the modules on replay-scored teacher data with SFT and then, for one branch, DPO. Step 3 (GRPO) continues from a Step 2 checkpoint and optimises against the replay reward directly, with the checkpoint/hint utility
U = mean_over_seeds( (y_hint - y_base) + 0.25 * (q_hint - q_base) ) in [-1.25, +1.25]
where y is the AppWorld verifier's pass/fail and q its partial-credit score.
What is in this repository
All twelve adapters are PEFT LoRA on the same base model and revision: Qwen/Qwen3-8B @ b968826d9c46dd6066d109eabc6255188de91218, r=16, alpha=32, dropout=0.05, no bias, targets q_proj,k_proj,v_proj,o_proj,gate_proj,up_proj,down_proj.
| subfolder | module | stage | part of evaluated system | size |
|---|---|---|---|---|
step2_supervision/selector_sft |
selector | Step 2 SFT | Factorized SFT, Factorized SFT + DPO decoder | 177.5 MiB |
step2_supervision/decoder_sft |
decoder | Step 2 SFT | Factorized SFT | 177.6 MiB |
step2_supervision/decoder_dpo |
decoder | Step 2 DPO | Factorized SFT + DPO decoder | 177.5 MiB |
step2_supervision/joint_sft |
joint | Step 2 SFT | Joint SFT | 177.6 MiB |
step2_supervision/joint_dpo |
joint | Step 2 DPO | Joint DPO | 177.6 MiB |
step3_grpo/R6-SFT/joint |
joint | Step 3 GRPO | GRPO R6-SFT (joint) | 177.5 MiB |
step3_grpo/R2-SFT/selector |
selector | Step 3 GRPO | GRPO R2-SFT (factorized, gate on) | 177.5 MiB |
step3_grpo/R2-SFT/decoder |
decoder | Step 3 GRPO | GRPO R2-SFT (factorized, gate on) | 177.5 MiB |
step3_grpo/R4-SFT/selector |
selector | Step 3 GRPO | GRPO R4-SFT (factorized, gate off) | 177.5 MiB |
step3_grpo/R4-SFT/decoder |
decoder | Step 3 GRPO | GRPO R4-SFT (factorized, gate off) | 177.5 MiB |
step3_grpo/R2-DPO/selector |
selector | Step 3 GRPO | GRPO R2-DPO (factorized, DPO decoder init) | 177.5 MiB |
step3_grpo/R2-DPO/decoder |
decoder | Step 3 GRPO | GRPO R2-DPO (factorized, DPO decoder init) | 177.5 MiB |
Each adapter subfolder holds adapter_config.json, adapter_model.safetensors, chat_template.jinja, tokenizer.json, tokenizer_config.json, a mix_provenance.json provenance record and its own README.md card. The Trainer's training_args.bin is deliberately not published: it is a torch pickle, it is not needed to load a LoRA adapter, and every field it holds is already in mix_provenance.json -> training.training_arguments as plain JSON. The tokenizer_config.json is byte-identical to the file written at training time, which is why it still carries three inert loader flags (local_files_only, is_local, backend); from_pretrained takes those from its own call arguments and ignores them.
Two evaluated systems in the tables below use no adapter at all: the factorized and joint zero-shot curators are the bare base model with the same prompts. They are the reference line.
results/ carries the small Step 3 artifacts: priority4_tables.json (all ten Dev summaries re-tabulated), deltas.json, training_runs.json, recompute_priority4_tables.py, and PRIORITY4_REPORT_ZH.md - the full Step 3 report, written in Chinese. Those five files are verbatim run records: they contain absolute paths, pod hostnames and job ids from the machines the campaign ran on, and the Chinese report's internal relative links point at files that are not part of this repository. recompute_priority4_tables.py is included for provenance only - it is stdlib-only, but it reads the private campaign root and cannot be re-run from this repository's contents. The same is true of build_upload_tree.py at the root: it is the script that staged this tree and recomputed every number in these cards, published so the derivation is auditable, and its path defaults point at the private campaign root.
MANIFEST.json at the root lists every staged file except itself - size, sha256, and which campaign artifact it came from - plus the adapter weight-hash cross-check and the recomputed macro table.
How these numbers were measured
- Dev pool: a frozen three-actor pool of 328 failed trajectories - Qwen3-32B 139, Qwen3.5-122B-A10B-FP8 129, Qwen3.8-27B 60. It is disjoint from the training manifest by scenario and by task. The three actors are the vLLM-served models
Qwen/Qwen3-32B,Qwen/Qwen3.5-122B-A10B-FP8andQwen/Qwen3.8-27B; in the raw summaries they are keyedqwen3_32b,qwen3_5_122b_a10bandqwen3_8_27b. - Protocol: one deterministic curator intervention per row - greedy generation for the decoder and joint modules, argmax over scored checkpoints for the selector (see How to load and run),
max_tokens=192. Each intervention is replayed under seeds 101/102/103 against a matched base branch. - Metrics:
Repair@3- fraction of all N rows where at least one of the three hint replays passes the verifier. Invalid curator outputs stay in N and count as failed repairs.Base@3- the same for the matched base branch, over the rows where all three base replays exist. Missing bases are never imputed.mean-pass@3- mean per-seed pass rate of the hint branch over all N rows.mean U- mean of the utility above over rows where all three seeds paired successfully; range [-1.25, +1.25].harmful- fraction of those paired rows with U < 0.
- Macro is the equal-weight mean over the three actors, not a pooled average, so the 60-row actor carries the same weight as the 139-row actor. Counts (
invalid,N) are sums. - Every number in this card was recomputed from the ten
summary.jsonfiles bybuild_upload_tree.py; the recomputed macro matched the value recorded in each summary to within 1.1e-16.
Dev results - equal-weight macro over three actors
| System | Repair@3 | Base@3 | mean-pass@3 | mean U | harmful | invalid | N |
|---|---|---|---|---|---|---|---|
| Factorized zero-shot (no adapter) | 17.1% | 21.8% | 11.1% | -0.018 | 34.1% | 0 | 328 |
| Factorized SFT | 33.6% | 13.1% | 29.2% | +0.232 | 16.4% | 0 | 328 |
| Factorized SFT + DPO decoder | 32.9% | 12.6% | 26.8% | +0.217 | 10.5% | 2 | 328 |
| Joint zero-shot (no adapter) | 19.6% | 22.0% | 12.0% | -0.011 | 34.5% | 0 | 328 |
| Joint SFT | 42.7% | 16.0% | 33.2% | +0.263 | 17.9% | 2 | 328 |
| Joint DPO | 35.4% | 15.2% | 27.1% | +0.194 | 14.3% | 1 | 328 |
| GRPO R6-SFT (joint) | 41.3% | 13.9% | 33.8% | +0.294 | 12.8% | 1 | 328 |
| GRPO R2-SFT (factorized, gate on) | 34.1% | 13.6% | 26.6% | +0.204 | 14.2% | 0 | 328 |
| GRPO R4-SFT (factorized, gate off) | 28.3% | 13.6% | 23.3% | +0.172 | 11.6% | 0 | 328 |
| GRPO R2-DPO (factorized, DPO decoder init) | 31.4% | 13.7% | 26.1% | +0.199 | 12.8% | 1 | 328 |
Dev results - per actor
| System | Actor | N | Repair@3 | Base@3 | mean-pass@3 | mean U | harmful | invalid |
|---|---|---|---|---|---|---|---|---|
| Factorized zero-shot | Qwen3-32B | 139 | 6.5% | 13.7% | 3.4% | -0.046 | 38.8% (54/139) | 0 |
| Factorized zero-shot | Qwen3.5-122B-A10B-FP8 | 129 | 14.7% | 20.2% | 6.5% | -0.028 | 43.4% (56/129) | 0 |
| Factorized zero-shot | Qwen3.8-27B | 60 | 30.0% | 31.7% | 23.3% | +0.019 | 20.0% (12/60) | 0 |
| Factorized SFT | Qwen3-32B | 139 | 11.5% | 10.1% | 7.9% | +0.036 | 22.3% (31/139) | 0 |
| Factorized SFT | Qwen3.5-122B-A10B-FP8 | 129 | 31.0% | 9.3% | 27.9% | +0.240 | 18.6% (24/129) | 0 |
| Factorized SFT | Qwen3.8-27B | 60 | 58.3% | 20.0% | 51.7% | +0.419 | 8.3% (5/60) | 0 |
| Factorized SFT + DPO decoder | Qwen3-32B | 139 | 12.2% | 8.6% | 7.7% | +0.044 | 16.8% (23/137) | 2 |
| Factorized SFT + DPO decoder | Qwen3.5-122B-A10B-FP8 | 129 | 34.9% | 9.3% | 31.5% | +0.287 | 13.2% (17/129) | 0 |
| Factorized SFT + DPO decoder | Qwen3.8-27B | 60 | 51.7% | 20.0% | 41.1% | +0.320 | 1.7% (1/60) | 0 |
| Joint zero-shot | Qwen3-32B | 139 | 11.5% | 15.1% | 6.5% | -0.023 | 41.0% (57/139) | 0 |
| Joint zero-shot | Qwen3.5-122B-A10B-FP8 | 129 | 15.5% | 22.5% | 6.7% | -0.032 | 42.6% (55/129) | 0 |
| Joint zero-shot | Qwen3.8-27B | 60 | 31.7% | 28.3% | 22.8% | +0.021 | 20.0% (12/60) | 0 |
| Joint SFT | Qwen3-32B | 139 | 16.5% | 11.7% | 10.1% | +0.051 | 24.1% (33/137) | 2 |
| Joint SFT | Qwen3.5-122B-A10B-FP8 | 129 | 46.5% | 16.3% | 35.1% | +0.292 | 14.7% (19/129) | 0 |
| Joint SFT | Qwen3.8-27B | 60 | 65.0% | 20.0% | 54.4% | +0.447 | 15.0% (9/60) | 0 |
| Joint DPO | Qwen3-32B | 139 | 14.4% | 10.9% | 7.2% | +0.023 | 23.2% (32/138) | 1 |
| Joint DPO | Qwen3.5-122B-A10B-FP8 | 129 | 43.4% | 13.2% | 35.7% | +0.312 | 13.2% (17/129) | 0 |
| Joint DPO | Qwen3.8-27B | 60 | 48.3% | 21.7% | 38.3% | +0.246 | 6.7% (4/60) | 0 |
| GRPO R6-SFT (joint) | Qwen3-32B | 139 | 15.8% | 12.3% | 8.9% | +0.047 | 21.0% (29/138) | 1 |
| GRPO R6-SFT (joint) | Qwen3.5-122B-A10B-FP8 | 129 | 46.5% | 9.3% | 37.5% | +0.351 | 12.4% (16/129) | 0 |
| GRPO R6-SFT (joint) | Qwen3.8-27B | 60 | 61.7% | 20.0% | 55.0% | +0.485 | 5.0% (3/60) | 0 |
| GRPO R2-SFT (factorized, gate on) | Qwen3-32B | 139 | 15.1% | 10.1% | 8.6% | +0.046 | 22.3% (31/139) | 0 |
| GRPO R2-SFT (factorized, gate on) | Qwen3.5-122B-A10B-FP8 | 129 | 35.7% | 10.9% | 32.3% | +0.294 | 7.0% (9/129) | 0 |
| GRPO R2-SFT (factorized, gate on) | Qwen3.8-27B | 60 | 51.7% | 20.0% | 38.9% | +0.271 | 13.3% (8/60) | 0 |
| GRPO R4-SFT (factorized, gate off) | Qwen3-32B | 139 | 15.8% | 10.1% | 7.7% | +0.037 | 23.7% (33/139) | 0 |
| GRPO R4-SFT (factorized, gate off) | Qwen3.5-122B-A10B-FP8 | 129 | 35.7% | 10.9% | 34.4% | +0.320 | 6.2% (8/129) | 0 |
| GRPO R4-SFT (factorized, gate off) | Qwen3.8-27B | 60 | 33.3% | 20.0% | 27.8% | +0.159 | 5.0% (3/60) | 0 |
| GRPO R2-DPO (factorized, DPO decoder init) | Qwen3-32B | 139 | 13.7% | 10.8% | 8.4% | +0.042 | 23.2% (32/138) | 1 |
| GRPO R2-DPO (factorized, DPO decoder init) | Qwen3.5-122B-A10B-FP8 | 129 | 35.7% | 8.5% | 32.6% | +0.296 | 10.1% (13/129) | 0 |
| GRPO R2-DPO (factorized, DPO decoder init) | Qwen3.8-27B | 60 | 45.0% | 21.7% | 37.2% | +0.261 | 5.0% (3/60) | 0 |
Each GRPO arm against its own initialization
| arm | initialization | arm Repair@3 | init Repair@3 | delta | arm mean U | init mean U | delta | arm harmful | init harmful |
|---|---|---|---|---|---|---|---|---|---|
| R6-SFT | Joint SFT | 41.3% | 42.7% | -1.4 pp | +0.294 | +0.263 | +0.031 | 12.8% | 17.9% |
| R2-SFT | Factorized SFT | 34.1% | 33.6% | +0.5 pp | +0.204 | +0.232 | -0.028 | 14.2% | 16.4% |
| R4-SFT | Factorized SFT | 28.3% | 33.6% | -5.3 pp | +0.172 | +0.232 | -0.060 | 11.6% | 16.4% |
| R2-DPO | Factorized SFT + DPO decoder | 31.4% | 32.9% | -1.5 pp | +0.199 | +0.217 | -0.018 | 12.8% | 10.5% |
What we found
- A trained curator is worth a lot over zero-shot. Every trained system beats its own zero-shot reference line by a wide margin on macro Repair@3 (19.6% joint zero-shot and 17.1% factorized zero-shot, versus 32.9%-42.7% for the four trained Step 2 systems), and the zero-shot curators sit at mean U around zero with a harmful rate above 34% - they damage recoverable checkpoints about as often as they help.
- Joint SFT has the highest macro Repair@3 measured here, but it is not uniformly the best system. Joint SFT reaches 42.7% macro Repair@3, +1.4 pp above the best GRPO arm (R6-SFT, 41.3%) - a gap inside the replay-noise band described under Limitations. On the other three headline metrics R6-SFT is ahead: mean U +0.294 vs +0.263, mean-pass@3 33.8% vs 33.2%, harmful rate 12.8% vs 17.9%. What is outside the band is the joint-versus-factorized gap: Joint SFT leads the best factorized Step 2 system (Factorized SFT, 33.6%) by +9.1 pp on macro Repair@3, and that lead survives GRPO.
- The interface is deliberately minimal. The curator emits only a checkpoint
step_idand a hint of at most two sentences - no taxonomy label, no structured diagnosis. Earlier ablations in the source repository picked this minimal output schema over richer ones; that comparison is prior work in the source repo and was not re-measured in the runs reported here, where every system already uses the minimal schema. - No GRPO arm clearly beat its own initialization on macro Repair@3. The four deltas are R6-SFT -1.4 pp, R2-SFT +0.5 pp, R4-SFT -5.3 pp, R2-DPO -1.5 pp. The best arm overall is R6-SFT (41.3%), still below its own initialization (42.7%). R6-SFT is the only arm that improved mean U (+0.031) while cutting the harmful rate, i.e. it made interventions safer rather than more often successful.
- The absolute-benefit gate helps.
R2-SFTandR4-SFTdiffer in exactly one variable - the gate that filters GRPO groups without a meaningful positive example. Gate on minus gate off is +5.9 pp macro Repair@3 and +0.032 mean U. Almost all of it comes from one actor (Qwen3.8-27B: +18.3 pp), which contributes only 60 rows, so treat the size of the gap with care. On the training side the gate-off arm did produce many more non-zero gradient batches; those extra updates did not pay off on Dev. - Curators pick late checkpoints. Over every intervention actually replayed during GRPO training, the chosen
step_idsits at a median of 0.857-0.900 of the trajectory length across the four arms (results/training_runs.json->runs.<arm>.scored_interventions.step_fraction_distribution). The repair is usually a late correction rather than a restart.
Limitations - please read before citing a number
- Single run, descriptive only. Every Dev number here comes from one evaluation pass with three replay seeds and one training seed (42). No significance testing was performed for the GRPO arms, and no confidence intervals are reported in this card.
- Replay is noisy. The following are values recorded by the Step 2 Dev report, not recomputed here (see
results/PRIORITY4_REPORT_ZH.mdsection 5.1): re-running the same curator over the same rows moved the two-actor macro Repair@3 by -0.8 to +4.5 percentage points, and by up to +7.8 on a single actor; when only the replay was repeated and the intervention held fixed, 2-14% of individual rows still flipped outcome. A paired bootstrap over the same 328 rows gave descriptive intervals roughly +/-5-6 percentage points wide. Differences of about +/-5 macro percentage points are inside that noise band and should not be read as results. - N = 328 rows is not 328 independent problems. The Dev pool covers only 55 distinct AppWorld tasks across 19 scenarios; the 60-row Qwen3.8-27B actor covers 30 tasks and 13 scenarios. Rows from the same task or scenario are correlated, so the effective sample is smaller than N suggests - which matters most for the per-actor numbers and for the gate comparison that rests on that smallest actor.
- An uncontrolled environment difference sits inside the arm-vs-initialization comparison. The Step 2 baselines and the Step 3 arms were evaluated on different days with different actor-server CUDA-graph modes. That mode is not recorded in any identity hash and cannot be separated out of the numbers.
- Actor-specific and AppWorld-only. A curator is trained against particular actor models and a particular benchmark. Nothing here shows transfer to another agent, another tool API, or another benchmark. The three Dev actors are the same family the curators were trained on.
- Base@3 moves with the curator. Different curators choose different checkpoints, so the matched base rate is not constant across rows; the paired
mean Uis the metric that controls for this, andBase@3should not be read as a fixed baseline. - These are LoRA adapters, not merged models. They only behave as measured on
Qwen/Qwen3-8Bat revisionb968826d9c46dd6066d109eabc6255188de91218with the curator prompts from the source repository. Loading them on another base or another revision is untested. - The Dev gate was released under a recorded human approval, not by rule. A small number of invalid curator outputs left paired coverage just below 1.0, so the automatic gate did not pass on its own rule; the release was recorded as a human decision with advisor sign-off still pending. Details are in
results/PRIORITY4_REPORT_ZH.md.
How to load and run
from transformers import AutoModelForCausalLM
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", revision="b968826d9c46dd6066d109eabc6255188de91218", dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(base, "Ameame1002/RLCurator", subfolder="step2_supervision/joint_sft")
model.eval()
For the factorized architecture load two adapters onto the same base and activate them by name:
from transformers import AutoModelForCausalLM
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", revision="b968826d9c46dd6066d109eabc6255188de91218", dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(base, "Ameame1002/RLCurator", subfolder="step2_supervision/selector_sft", adapter_name="selector")
model.load_adapter("Ameame1002/RLCurator", subfolder="step2_supervision/decoder_sft", adapter_name="decoder")
model.set_adapter("selector") # -> step_id
model.set_adapter("decoder") # -> repair_hint
The adapters expect the curator prompts from the source repository (prompts/factorized_curator/{selector,decoder,joint}_v1.txt at https://github.com/22856226/RlCurator), and the tokenizer/chat template shipped in each subfolder is the one used at training and evaluation time.
The two module types are not decoded the same way. The decoder and the joint module generate: greedy, max_new_tokens=192, thinking disabled (enable_thinking=False when applying the chat template). The selector does not generate - it scores the one canonical JSON string {"step_id": i} per legal checkpoint i and takes the argmax of those canonical-JSON+EOS sequence log-probabilities, ties going to the earliest step. The softmax over the same scores is recorded alongside the choice as a distribution (temperature 1, epsilon=0) but it is not sampled from at evaluation time. That is the protocol recorded in every Dev summary (protocol.selector_decoding, protocol.categorical_distribution, protocol.selector_tie_break).
The snippets above use dtype=, which needs transformers >= 4.56; on older versions pass torch_dtype="bfloat16" instead. The recorded training and evaluation environment was transformers 5.16.1, peft 0.20.0, torch 2.13.0, Python 3.11.
Reproducing the reward or the Dev numbers additionally requires an AppWorld installation with the checkpoint snapshots, which are not part of this repository.
Provenance and integrity
Every adapter ships a mix_provenance.json recording the base model and revision, the parent adapter it was initialized from, the training manifest hash, and the sha256 of its own weight file. MANIFEST.json at the repository root repeats the sha256 of every staged file except itself (a manifest cannot contain its own hash) and records, per adapter, that the staged adapter_model.safetensors hashes to exactly the adapter_weight_sha256 written in its provenance at training time.
One class of change was applied when copying the Step 3 report artifacts: real email addresses were replaced with [redacted-email] - 2 occurrences in results/PRIORITY4_REPORT_ZH.md, 1 occurrence in results/priority4_tables.json. All of them are the same personal address, recorded as the approver of the Dev coverage gate. Addresses on RFC 2606 reserved domains (example.com and friends) are placeholders, not mailboxes, and are left verbatim - which is why a quoted git-author string in the Chinese report still reads Andy <andy@example.com>. Service and log file names that merely contain an @ are untouched. MANIFEST.json records the source sha256 and the staged sha256 for every redacted file, and the remaining results/ artifacts are byte-identical to their sources. Nothing else was changed.
Citation
A paper on CuratorRL is in preparation. Until it is public, please cite this artifact repository:
@misc{curatorrl,
title = {CuratorRL: Replay-Validated Trajectory Repair for Tool-Using Agents},
author = {CuratorRL authors},
year = {2026},
note = {Paper in preparation. Model artifacts: https://huggingface.co/Ameame1002/RLCurator},
url = {https://github.com/22856226/RlCurator}
}
License
Apache-2.0 for the adapter weights and the material in this repository. Qwen/Qwen3-8B remains under its own license, and AppWorld under its own; neither is redistributed here.
中文摘要
CuratorRL 训练一个 curator(修复器):读入一条失败的 AppWorld 工具调用轨迹,输出一个最小修复——从哪个 checkpoint(step_id)重来,以及一句不超过两句话的 repair_hint。关键在于这个输出怎么被验证:不是因为老师模型写了它就算数,而是真的把环境回滚到那个 checkpoint、重放一遍,并与同一 checkpoint、同一 seed 的 base 分支对照,两边都交给 AppWorld 的 ground-truth verifier 判分。
本仓库是 12 个 LoRA adapter:Step 2 监督阶段 5 个(factorized 的 selector/decoder,以及 joint),Step 3 GRPO 阶段 7 个(4 个臂)。上面的 Dev 结果表是三个 actor 的等权宏平均,328 行,单次运行、三个回放 seed、未做显著性检验;无效输出留在分母里记为修复失败,缺失的 base 不插补。完整的 Step 3 报告(中文)见 results/PRIORITY4_REPORT_ZH.md;该报告里指向仓库其他文件的相对链接在本仓库中不会解析。
- Downloads last month
- -