Publications · works cited, in full

Publications

12 papers, 2024–2026: 9 published, 3 preprints. First or (co-)first author on 10.

Index of subjects

153113611

Show Order

2026

  1. Lost in Serialization: Invariance and Generalization of LLM Graph Reasoners

    D. Herbst*, L. Karbevska*, D. Kumar*, A. Ahuja*, F. G. Nasrabadi*, F. Frasca

    OpenReview

    While promising, graph reasoners based on Large Language Models (LLMs) lack built-in invariance to symmetries in graph representations. Operating on sequential graph serializations, LLMs can produce different outputs under node reindexing, edge reordering, or formatting changes, raising robustness concerns. We systematically analyze these effects, studying how fine-tuning impacts encoding sensitivity as well generalization on unseen tasks. We propose a principled decomposition of graph serializations into node labeling, computational structure, and surface encoding, and evaluate LLM robustness to variations of each of these factors on a comprehensive benchmarking suite. We also contribute a novel set of spectral tasks to further assess generalization abilities of fine-tuned reasoners. Results show that larger (non-fine-tuned) models are more robust, and fine-tuning reduces sensitivity to node relabeling but may increase it to variations in structure and format, while it does not consistently improve performance on unseen tasks.

    Graph Reasoning · LLM · NLP

    GCLR @ AAAI 2026; GFM @ ICML 2026 Jul 2026 · poster co-first
  2. Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

    D. Kumar, N. A. Birur, T. Baswa, S. Agarwal, P. Harshangi

    Agentic systems are moving from prototypes to production, where they read untrusted inputs, call tools with real permissions, persist state, and act on a user's behalf, which expands the security and safety surface beyond chat-only models. Yet most standard evaluations remain single turn, for example multitask multiple choice and docstring-to-code tests, which are weak indicators of reliability in long-horizon settings where actions have consequences. Emerging agent benchmarks further show that interactive environments reveal qualitatively different failure modes compared with static leaderboards, underscoring the need for evaluations that track behavior over multi-step tasks and tool use. We present a systematic black-box framework for risk-aware agent evaluation that requires only basic system descriptions to initiate comprehensive red teaming. Our approach introduces: (1) a seven-domain taxonomy mapping observable behaviors to specific risk categories, (2) fully automated SAGE-RT powered red teaming producing 120 adversarial scenarios per domain without human intervention, and (3) human-in-the-loop evaluation of results where automated LLM judges provide initial scoring followed by expert validation of critical findings. Empirical validation across two production-ready agent architectures (single-agent CrewAI and multi-agent AutoGen) with four base models reveals alarming vulnerability patterns: 56.25% average governance risk across all systems, 65% privacy risk in multi-agent configurations, and critical agent behavior vulnerabilities reaching 85% in specific model-architecture combinations. Notably, our black-box approach discovered 98% of vulnerabilities identified by Unit 42's white-box analysis, while requiring 10x less manual effort. These findings demonstrate that systematic, taxonomy-guided evaluation can effectively identify architectural vulnerabilities without privileged access, providing a scalable path toward safer agent deployments.

    Agentic AI · Red Teaming · NLP

    LAMaS Workshop @ AAAI 2026 Jan 2026
  3. SocioEval: A Template-Based Framework for Evaluating Socioeconomic Status Bias in Foundation Models

    D. Kumar, I. Gupta, N. A. Birur, T. Baswa, S. Agarwal, P. Harshangi

    As Large Language Models (LLMs) increasingly power decision-making systems across critical domains, understanding and mitigating their biases becomes essential for responsible AI deployment. Although bias assessment frameworks have proliferated for attributes such as race and gender, socioeconomic status bias remains significantly underexplored despite its widespread implications in the real world. We introduce SocioEval, a template-based framework for systematically evaluating socioeconomic bias in foundation models through decision-making tasks. Our hierarchical framework encompasses 8 themes and 18 topics, generating 240 prompts across 6 class-pair combinations. We evaluated 13 frontier LLMs on 3,120 responses using a rigorous three-stage annotation protocol, revealing substantial variation in bias rates (0.42%-33.75%). Our findings demonstrate that bias manifests differently across themes lifestyle judgments show 10x higher bias than education-related decisions and that deployment safeguards effectively prevent explicit discrimination but show brittleness to domain-specific stereotypes. SocioEval provides a scalable, extensible foundation for auditing class-based bias in language models.

    Socioeconomic Status · Foundation Models · NLP

    3rd Workshop on AI Governance (AIGOV) at AAAI 2026 Jan 2026

2025

  1. Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations

    D. Kumar*, S. Jena*, N. A. Birur, T. Baswa, S. Agarwal, P. Harshangi

    OpenReview

    Multimodal large language models (MLLMs) have achieved remarkable progress, yet remain critically vulnerable to adversarial attacks that exploit weaknesses in cross-modal processing. We present a systematic study of multimodal jailbreaks targeting both vision-language and audio-language models, showing that even simple perceptual transformations can reliably bypass state-of-the-art safety filters. Our evaluation spans 1,900 adversarial prompts across three high-risk safety categories harmful content, CBRN (Chemical, Biological, Radiological, Nuclear), and CSEM (Child Sexual Exploitation Material) tested against seven frontier models. We explore the effectiveness of attack techniques on MLLMs, including FigStep-Pro (visual keyword decomposition), Intelligent Masking (semantic obfuscation), and audio perturbations (Wave-Echo, Wave-Pitch, Wave-Speed). The results reveal severe vulnerabilities: models with almost perfect text-only safety (0% ASR) suffer >75% attack success under perceptually modified inputs, with FigStep-Pro achieving up to 89% ASR in Llama-4 variants. Audio-based attacks further uncover provider-specific weaknesses, with even basic modality transfer yielding 25% ASR for technical queries. These findings expose a critical gap between text-centric alignment and multimodal threats, demonstrating that current safeguards fail to generalize across cross-modal attacks. The accessibility of these attacks, which require minimal technical expertise, suggests that robust multimodal AI safety will require a paradigm shift toward boarder semantic-level reasoning to mitigate possible risks.

    Multimodal · Jailbreaking · Vision-Language · Audio · NLP

    NeurIPS Reliable ML from Unreliable Data Workshop Oct 2025 co-first
  2. Quantifying CBRN Risk in Frontier Models

    D. Kumar, N. A. Birur, T. Baswa, S. Agarwal, P. Harshangi

    Frontier Large Language Models (LLMs) pose unprecedented dual-use risks through the potential proliferation of chemical, biological, radiological, and nuclear (CBRN) weapons knowledge. We present the first comprehensive evaluation of 10 leading commercial LLMs against both a novel 200-prompt CBRN dataset and a 180-prompt subset of the FORTRESS benchmark, using a rigorous three-tier attack methodology. Our findings expose critical safety vulnerabilities: Deep Inception attacks achieve 86.0% success versus 33.8% for direct requests, demonstrating superficial filtering mechanisms; Model safety performance varies dramatically from 2% (claude-opus-4) to 96% (mistral-small-latest) attack success rates; and eight models exceed 70% vulnerability when asked to enhance dangerous material properties. We identify fundamental brittleness in current safety alignment, where simple prompt engineering techniques bypass safeguards for dangerous CBRN information. These results challenge industry safety claims and highlight urgent needs for standardized evaluation frameworks, transparent safety metrics, and more robust alignment techniques to mitigate catastrophic misuse risks while preserving beneficial capabilities.

    CBRN · Risk · Frontier Models · NLP

    NeurIPS Reliable ML from Unreliable Data Workshop Oct 2025
  3. Beyond Western Politics: Cross-Cultural Benchmarks for Evaluating Partisan Associations in LLMs

    D. Kumar*, I. Gupta*, N. A. Birur, T. Baswa, S. Agarwal, P. Harshangi

    Partisan bias in LLMs has been evaluated to assess political leanings, typically through a broad lens and largely in Western contexts. We move beyond identifying general leanings to examine harmful, adversarial representational associations around political leaders and parties. To do so, we create datasets NeutQA-440 (non-adversarial prompts) and AdverQA-440 (adversarial prompts), which probe models for comparative plausibility judgments across the USA and India. Results show high susceptibility to biased partisan associations and pronounced asymmetries (e.g., substantially more favorable associations for U.S. Democrats than Republicans) alongside mixed-polarity concentration around India's BJP, highlighting systemic risks and motivating standardized, cross-cultural evaluation.

    Partisan Bias · Large Language Models · NLP

    NeurIPS LLM Evaluation Workshop Sep 2025 · poster co-first
  4. No Free Lunch with Guardrails

    D. Kumar, N. A. Birur, T. Baswa, S. Agarwal, P. Harshangi

    arXiv

    As large language models (LLMs) and generative AI become widely adopted, guardrails have emerged as a key tool to ensure their safe use. However, adding guardrails isn't without tradeoffs; stronger security measures can reduce usability, while more flexible systems may leave gaps for adversarial attacks. In this work, we explore whether current guardrails effectively prevent misuse while maintaining practical utility. We introduce a framework to evaluate these tradeoffs, measuring how different guardrails balance risk, security, and usability, and build an efficient guardrail. Our findings confirm that there is no free lunch with guardrails; strengthening security often comes at the cost of usability. To address this, we propose a blueprint for designing better guardrails that minimize risk while maintaining usability.

    Guardrails · AI Safety · Security · Usability · LLMs

    arXiv preprint Apr 2025 preprint

2024

  1. Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs

    D. Kumar*, U. Jain*, S. Agarwal, P. Harshangi

    OpenReview

    We conduct a comprehensive investigation of implicit bias in large language models through systematic evaluation of over 50 different LLMs. Our study reveals concerning patterns of bias that persist across different model architectures, training methodologies, and deployment strategies. We provide insights into the sources of these biases and propose mitigation strategies for safer AI deployment.

    Bias · Fairness · LLMs · Safety · Evaluation

    NeurIPS Safe Generative AI Workshop Dec 2024 · poster co-first
  2. SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming

    A. Kumar*, D. Kumar*, J. Loya, N. A. Birur, T. Baswa, S. Agarwal, P. Harshangi

    OpenReview

    We introduce SAGE-RT, a comprehensive framework for generating synthetic alignment data tailored for safety evaluation and red teaming of generative AI systems. Our approach addresses the critical need for diverse, high-quality evaluation datasets in AI safety research by providing automated generation of challenging test cases.

    Red Teaming · AI Safety · Synthetic Data · Alignment · Evaluation

    NeurIPS Red Teaming GenAI Workshop Oct 2024 · poster co-first
  3. Efficacy of the SAGE-RT Dataset for Model Safety Alignment: A Comparative Study

    T. Baswa, N. A. Birur, D. Kumar, J. Loya, A. Kumar, P. Harshangi, S. Agarwal

    OpenReview

    This work presents a thorough evaluation of the SAGE-RT dataset's impact on model safety alignment. Through comparative studies across various model architectures and safety metrics, we demonstrate the dataset's effectiveness in improving alignment while maintaining model utility.

    Safety Alignment · Dataset Evaluation · Model Safety · Comparative Study

    NeurIPS Pluralistic Alignment Workshop Oct 2024 · poster
  4. VERA: Validation and Enhancement for Retrieval Augmented Systems

    N. A. Birur, T. Baswa, D. Kumar, J. Loya, S. Agarwal, P. Harshangi

    arXiv

    We present VERA, a comprehensive framework for validation and enhancement of retrieval-augmented generation systems. Our approach addresses critical challenges in RAG system deployment including retrieval quality, generation consistency, and system reliability through systematic validation and enhancement techniques.

    RAG · Retrieval Systems · Validation · Enhancement · NLP

    arXiv preprint Sep 2024 preprint
  5. Increased LLM Vulnerabilities from Fine-tuning and Quantization

    D. Kumar, A. Kumar, S. Agarwal, P. Harshangi

    arXiv

    Model quantization and fine-tuning have become essential techniques for deploying large language models in resource-constrained environments. However, their impact on model security has been largely overlooked. In this work, we conduct the first systematic study of how quantization and fine-tuning affect adversarial robustness in LLMs. Through comprehensive experiments across multiple model families and quantization schemes, we demonstrate that aggressive quantization can increase attack success rates significantly compared to full-precision models.

    Quantization · Adversarial Attacks · Model Compression · Security · LLMs

    arXiv preprint Apr 2024 preprint