PolicyShiftGuard-7B-RP-SFT

\ud83d\udcc3 Paper | \ud83c\udf10 Project Page | \ud83d\udcbb GitHub

This repository releases the Stage-1 Randomized Policy SFT (RP-SFT) checkpoint for the 7B PolicyShiftGuard model.

RP-SFT is the first training stage in PolicyShiftGuard. It trains a Qwen2.5-VL guardrail model to read policy bundles under randomized policy identifiers and randomized policy ordering. This checkpoint is provided for reproducibility and ablation use. The final public model after the second-stage adaptation is available at PolicyShiftGuard/PolicyShiftGuard-7B.

Intended Use

Use this checkpoint when you want to reproduce the two-stage training pipeline or compare Stage-1 RP-SFT against the final BP-Adapt model.

For standard evaluation or deployment, use the final model instead:

Dataset

The model is trained with PolicyShiftBench supervision:

Notes

  • This is an intermediate checkpoint, not the final model reported as the main PolicyShiftGuard model.
  • This checkpoint corresponds to the randomized-policy no-think Stage-1 SFT setting.
  • Training-state files such as optimizer states are intentionally not included.

Citation

If you find this work helpful, please cite the paper:

@article{song2026policyshiftguard,
  title   = {PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails},
  author  = {Song, Mingyang and Xu, Luxin and Sun, Haoyu and Pan, Minzhou and Cheng, Yu and Li, Bo},
  journal = {arXiv preprint arXiv:2607.05910},
  year    = {2026}
}
Downloads last month
17
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PolicyShiftGuard/PolicyShiftGuard-7B-RP-SFT

Finetuned
(1140)
this model

Dataset used to train PolicyShiftGuard/PolicyShiftGuard-7B-RP-SFT

Paper for PolicyShiftGuard/PolicyShiftGuard-7B-RP-SFT