AtMem-MJM v1 pilot

AtMem Memory Judgment Model (AtMem-MJM) is a small LoRA adaptation of the 0.6B Skywork Reward V2 Qwen3 sequence-classification model. It assigns an uncalibrated scalar score to a task and one candidate memory. Higher scores mean the model ranked that candidate as more useful in the training objective; they are not probabilities or truth estimates.

AtMem-MJM only ranks candidates. Your application must decide which memories are eligible before scoring and must validate selected records afterward. Do not use this model to decide access, scope, freshness, sensitivity, deletion, consent, or policy.

  • Model weights: adapter only (about 2.3 MB); the pinned public base model is downloaded on first load.
  • Leaderboard and aggregate benchmark: AtMem-MJM leaderboard.
  • License: Apache-2.0 adapter; the base checkpoint is also Apache-2.0. See the base model card.
  • Intended use: research and local ranking experiments over already authorized candidate memories.

Quick start

Requires Python 3.10+ and PyTorch, Transformers, and PEFT. The first run downloads the adapter and the pinned ~0.6B base checkpoint from Hugging Face. Choose CPU, CUDA, or Apple MPS based on the installed PyTorch build.

pip install 'torch>=2.7' 'transformers>=5.0' 'peft>=0.17' huggingface_hub
hf download atmem/atmem-mjm-v1 --local-dir ./atmem-mjm-v1
cd atmem-mjm-v1

Save the following as rank.py in that folder. The adapter, tokenizer, and inference.py helper are included:

from inference import MemoryJudge

judge = MemoryJudge(model_id=".")  # first load fetches the pinned base model
candidates = [
    {"id": "memory-1", "text": "The user confirmed the new address yesterday."},
    {"id": "memory-2", "text": "An address appeared in a note from three years ago."},
]
for item in judge.rank("What is the customer's current address?", candidates):
    print(item["id"], item["score"])

The helper batches scoring, uses the model's Qwen chat template, truncates inputs at 512 tokens, and returns a stable descending order. For direct Hub loading without downloading the adapter folder first, use MemoryJudge() from a local script. Scores are meaningful only for relative ordering within a task; do not compare their magnitudes across unrelated tasks.

Evaluation

The pilot used the same 1,986 LoCoMo questions and frozen top-ten candidate pool for each reranking arm. Results are evidence-ranking metrics, not generated-answer accuracy:

Metric AtMem native order Jev 1.13.0 Skywork 0.6B, unmodified AtMem-MJM v1
MRR@5 0.4259 0.5880 0.4540 0.4632
Recall@1 0.3399 0.5433 0.3661 0.3741
Recall@10 0.6495 0.6495 0.6495 0.6495

AtMem-MJM modestly improved MRR@5 over unmodified Skywork in this exploratory run, but did not outperform Jev. The leaderboard reports paired conversation-cluster bootstrap intervals. There were only ten conversation clusters and four historical source evidence-reference anomalies remain in the primary denominator, so uncertainty is substantial.

Training and limitations

The adapter was trained from Skywork revision 8c14a4e9e6321deaf572544339b16b8d6bbe8886 using a Bradley–Terry pairwise objective. The pilot used 72 AtMem-authored synthetic training pairs, 24 family-held-out validation pairs, and 24 audit pairs from two untouched template families. Pair accuracy was 1.00 on each 24-pair split; these small, templated splits do not establish broad generalization. Training used rank-4 LoRA on q_proj and v_proj, with the reward score head saved; three epochs ran in float32 on Apple MPS.

The adapter is not an agent, memory store, retrieval system, or safety policy. The experiment does not establish generated-answer quality, real-world memory quality, authorization correctness, tool-use safety, crash continuity, or exactly-once behavior. Human semantic review of broader preference data and evaluation on more tasks are still needed.

The synthetic training dataset remains in a separate private repository and is not needed to run inference. No benchmark questions, candidate memories, evidence identifiers, private histories, or evaluation score rows are included in this model repository.

Citation

@misc{taghia2026atmemmjm,
  author = {Taghia, Javad and AtMem contributors},
  title = {AtMem Memory Judgment Model (AtMem-MJM) v1 pilot},
  year = {2026},
  url = {https://huggingface.co/atmem/atmem-mjm-v1}
}
Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for atmem/atmem-mjm-v1

Finetuned
Qwen/Qwen3-0.6B
Adapter
(1)
this model