Papers

AI safety, interpretability and evaluation research, most of it co-authored with the early-career researchers I mentor at Algoverse and PRISM. Every submission, review and decision is public on OpenReview.

7 papers accepted in 2026 ICML · ICLR · KDD · EACL · WACV My OpenReview profile →
2026

Mutual Information Transfer Regularization for Logical Consistency: A Controlled Empirical Study

A binary QA model can answer a question correctly and still give the same label to its negation. We test whether Mutual Information Transfer Regularization, which discourages adjacent transformer layers from making redundant residual contributions, reduces those contradictions. Across an estimator screen, three backbones, a weight sweep and a five-seed replication, the answer is a clean negative: the gains are small, unstable, and not competitive with direct supervision on negated examples.

OpenReview

Preference Optimization Drives Monoculture in LLM Prediction Markets

Prediction markets aggregate information only because traders hold independent beliefs. We show that preference-optimised LLM traders lose that independence: alignment training pulls them toward a shared prior, and the market collapses into a correlated monoculture that prices confidently and wrongly.

OpenReview arXiv

Interpreting Latent CoT Reasoning as Dynamical Systems

Latent reasoning methods like CODI and COCONUT carry several superimposed candidate traces in hidden space, which makes their reasoning hard to read. We model the latent token sequence as a trajectory in representation space and analyse it as a dynamical system. Latent CoT turns out to have structured, non-random dynamics with two stability classes: CODI behaves as a stable attractor, COCONUT as an unstable expanding system.

OpenReview arXiv

Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents

LLMs are increasingly proposed as runtime oversight for autonomous agents, but nobody audits the explainer itself. In a black-box audit of an LLM-augmented active inference agent on German energy grid data, 0 of 30 explanations flagged an adversarial 600 MW observation injection, and all three backends rationalised objectively wrong actions 80–95% of the time. Fluent explanation and faithful oversight are different capabilities, and the explainer pattern quietly conflates them.

OpenReview

Cross-Lingual Fairness Drift in LLM Moral Reasoning

A framework for measuring how LLM ethical reasoning shifts across linguistic subpopulations. Five models, 50 moral dilemmas, four languages (English, Spanish, Korean, Mandarin) and a seven-pillar rubric: consequentialist bias amplifies outside English and cultural grounding collapses by up to 88%. Language turns out to be a leading indicator of behavioural drift risk in deployed systems.

OpenReview

FluffInjector: Diagnosing Logical Consistency Failures in Chain-of-Thought Reward Models

We identify fluff injection: replacing a logically necessary reasoning step with plausible-sounding commentary. Our benchmark pairs each problem with a clean chain and a chain where 25–40% of steps are non-inferential filler, same length, same answer. Frontier judges validate the fluffed chains far too often. The verifier we train on it, SmartRM, cuts false positives from 37.4% to 2.7% at 97.3% overall accuracy.

OpenReview

Quantifying Text Dominance in Multimodal AI

Vision-language models are being proposed as scalable tools for judging “substantial similarity” in copyright disputes. Our adversarial “Gaslight Protocol” shows they follow the prompt over the pixels: primed with an authoritative legal accusation, frontier models assign high similarity scores to visually unrelated images. We define a Blindness Index to quantify it, and argue these models are unfit for automated copyright enforcement without human verification.

OpenReview

HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance

Tool-retrieval benchmarks tend to use clean, well-formed queries that nobody actually types. HumanMCP is a dataset of human-like queries for measuring whether agents retrieve the right Model Context Protocol tool under realistic phrasing.

arXiv
2025
Earlier

Earlier applied ML work

Compositional Attention Networks for interpretability in natural language question answering; CNNs for classifying cancer through DNA methylation; detecting parking spaces in a parcel from satellite imagery; and TinyML for solar panels, bringing edge computing to solar energy systems.

Google Scholar