PolicyShiftGuard-7B / README.md
hitsmy's picture
Add pipeline tag, library name, and project links to model card (#1)
e5e41ed
|
Raw
History Blame Contribute Delete
1.93 kB
metadata
base_model: Qwen/Qwen2.5-VL-7B-Instruct
datasets:
  - PolicyShiftBench/PolicyShiftBench
license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
tags:
  - vision-language
  - image-safety
  - guardrails
  - policy-conditioned
  - qwen2.5-vl

PolicyShiftGuard-7B

📜 Paper | 💻 Code | 🏠 Project Page

PolicyShiftGuard-7B is a policy-conditioned image guardrail model based on Qwen2.5-VL-7B. It is trained to follow a supplied policy bundle and produce structured image-safety decisions under changing application policies.

Expected Output Format

true | <two-digit risk category id> | <short reason>
false | <short reason>

Training Data

This checkpoint is trained with the PolicyShiftBench public data release:

  • Dataset: PolicyShiftBench/PolicyShiftBench
  • Main evaluation splits: ID/adaptive branch and OOD/shift branch
  • Training stages: randomized policy SFT followed by boundary-pair policy adaptation

Intended Use

Use this model for research on policy-conditioned multimodal safety, adaptive image moderation, and robustness under policy shifts. The model should be evaluated with explicit policy bundles rather than as a fixed universal safety classifier.

Limitations

This is a research checkpoint. It may fail under policies, languages, visual domains, or deployment settings not represented in the benchmark. Outputs should not be treated as legal or compliance advice.

Citation

If you use this model, please cite the paper:

@article{song2026policyshiftguard,
  title   = {PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails},
  author  = {Song, Mingyang and Xu, Luxin and Sun, Haoyu and Pan, Minzhou and Cheng, Yu and Li, Bo},
  journal = {arXiv preprint arXiv:2607.05910},
  year    = {2026}
}