Temporal Dynamics of Safety in Diffusion-Based Language Models
This research project models, measures, and mitigates the temporal evolution of safety alignment within diffusion-based language models such as Dream-7B, TraDo-8B, LLaDA2.0 and many more. Unlike autoregressive models that generate tokens sequentially, diffusion models produce text via iterative denoising, creating unique safety vulnerabilities. The project aims to establish the first temporal safety framework for diffusion-based language models by quantifying when and how alignment fails during the generation process.
Questions
- How does the probability of unsafe token emergence evolve across denoising steps?
- Can step-wise or in-loop alignment reduce harmful emergence earlier than end-only filtering?
- Do jailbreaks and alignment failures transfer between diffusion and autoregressive models?
How
- Mechanistic Mapping
- Step-wise logging of denoising states with metrics like First Harmful Step (FHS), Irreversibility Index, and KL Drift
- Temporal Alignment
- In-loop guardrails including step-wise risk scoring, mask-aware gating, and learned temporal policy heads
- Cross-Architecture Transfer
- Applying DIJA and PAD jailbreaks across Dream-7B and autoregressive baselines
What comes out of it
- toolkitDiffusion Safety Probe: open toolkit for per-step risk visualization.
- papersAlignment Drift in Diffusion LMs; Temporal Alignment for Diffusion Decoders; Transferable Jailbreaks Across Architectures.
All planned; nothing here is published yet.
collaborationIf you work on diffusion language models, or on guardrails that have to run inside a decoder, I would like to compare notes. Write to me