Projects · in progress and finished

Projects

One question on my desk now, and the things I built before it.

§ 1now

Temporal Dynamics of Safety in Diffusion-Based Language Models

This research project models, measures, and mitigates the temporal evolution of safety alignment within diffusion-based language models such as Dream-7B, TraDo-8B, LLaDA2.0 and many more. Unlike autoregressive models that generate tokens sequentially, diffusion models produce text via iterative denoising, creating unique safety vulnerabilities. The project aims to establish the first temporal safety framework for diffusion-based language models by quantifying when and how alignment fails during the generation process.

Questions

  1. How does the probability of unsafe token emergence evolve across denoising steps?
  2. Can step-wise or in-loop alignment reduce harmful emergence earlier than end-only filtering?
  3. Do jailbreaks and alignment failures transfer between diffusion and autoregressive models?

How

Mechanistic Mapping
Step-wise logging of denoising states with metrics like First Harmful Step (FHS), Irreversibility Index, and KL Drift
Temporal Alignment
In-loop guardrails including step-wise risk scoring, mask-aware gating, and learned temporal policy heads
Cross-Architecture Transfer
Applying DIJA and PAD jailbreaks across Dream-7B and autoregressive baselines
thresholdfirst harmful stepstays harmfulrecoversT0denoising step
Fig.Schematic, not data. Risk of the partial output at each denoising step, from pure noise (T) to the finished text (0). The first harmful step is where it first crosses the line; irreversibility asks whether it ever comes back under.

What comes out of it

  • toolkitDiffusion Safety Probe: open toolkit for per-step risk visualization.
  • papersAlignment Drift in Diffusion LMs; Temporal Alignment for Diffusion Decoders; Transferable Jailbreaks Across Architectures.

All planned; nothing here is published yet.

collaborationIf you work on diffusion language models, or on guardrails that have to run inside a decoder, I would like to compare notes. Write to me

§ 2before, and alongside

Other projects

  1. SpeakerStream

    SpeakerStream aims to provide an accurate and efficient solution for speaker diarization and transcription from video sources in real-time.

    SpeakerStream is a comprehensive solution for real-time speaker diarization and transcription from video sources. The system combines advanced audio processing techniques with machine learning models to accurately identify different speakers and transcribe their speech in real-time applications.

    What it does

    • Real-time speaker diarization with high accuracy
    • Efficient video processing pipeline
    • Scalable architecture for multiple concurrent streams
    • Integration with popular streaming platforms

    Independent project.

    finished PythonPyTorchAudio ProcessingReal-time SystemsDocker
  2. EnvisEdge

    Edge computing solution for computer vision applications with optimized inference and deployment capabilities.

    EnvisEdge is an edge computing platform specifically designed for computer vision applications. The project focuses on optimizing deep learning models for deployment on edge devices while maintaining high accuracy and low latency for real-world applications.

    What I built

    • Optimized model deployment for edge devices
    • Real-time computer vision processing
    • Efficient resource utilization
    • Scalable edge computing architecture
    finished PythonTensorFlowEdge ComputingComputer VisionOptimization
  3. CommonLit Readability Prize

    ML models for rating the complexity of reading passages for Grade 3-12 classroom use, achieving competitive performance in Kaggle competition.

    This Kaggle competition project focused on developing machine learning models to automatically rate the complexity of reading passages for educational use in grades 3-12. The challenge involved creating models that could accurately assess text difficulty to help educators select appropriate reading materials for their students.

    What it does

    • Competitive ranking in Kaggle competition
    • RMSE-based evaluation for text complexity
    • Educational impact for classroom applications
    • Robust model performance across diverse text types
    finished PythonScikit-learnNLPFeature EngineeringKaggle
  4. VERA: Validation and Enhancement for RAG Systems

    Framework for validating and enhancing retrieval-augmented generation systems, addressing key challenges in RAG system reliability.

    VERA is a comprehensive framework designed to validate and enhance retrieval-augmented generation (RAG) systems. The project addresses critical challenges in RAG deployment including retrieval quality, generation consistency, and overall system reliability through systematic validation and enhancement techniques.

    In progress

    • RAG system validation methodologies
    • Enhancement techniques for retrieval quality
    • Consistency metrics for generation
    • Reliability frameworks for production deployment
    in progress PythonLangChainVector DatabasesRAGEvaluation