Skip to projects

Projects

Selected research in LLM alignment, privacy, and efficient inference.

AIMES MARS CSSD PROPS C-SQS CURATE Learning to Diagnose Privately STAMP

Adaptive Multi-Value Control in LLMs via Causal Activation Steering

LLM alignment · Activation steering · Multi-value control

AIMES feedback loop connecting activation extraction, value readouts, adaptive steering weights, and response generation.

LLMs often need to respond to several human values at once, and fixed steering strengths cannot adapt as a response develops. AIMES uses intermediate-layer readouts to monitor value-specific states and adjusts multiple activation interventions at each decoding step. Across model families and value combinations, it shows depth-dependent advantages over fixed joint steering while using smaller activation interventions and maintaining comparable response quality.

MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling

Reward modeling · Preference learning · Data augmentation

MARS allocates augmentation by reward margin, then refines semantically similar chosen and rejected responses.

Reward models must often learn from limited human preference data. MARS allocates more augmentation to preference pairs with small reward margins and uses semantic distance to identify pairs that need refinement before paraphrasing. Evaluations across three datasets and two reward-model backbones show improvements in average RewardBench performance and downstream alignment win rates over the evaluated augmentation baselines.

Conformal Speculative Speculative Decoding (CSSD)

Efficient inference · Conformal prediction · Speculative decoding

Accepted for oral presentation at Asilomar 2026. Full paper forthcoming.

SSD uses geometric fanout for speculative caching; CSSD uses context-dependent hazard estimates and conformal calibration to select cache entries.

Speculative speculative decoding precomputes continuations that can be reused after target-model verification. CSSD uses context-dependent rejection modeling to estimate plausible accepted lengths and online conformal calibration to select which lengths to retain in the cache. Cached continuations are used when verification produces a matching outcome; otherwise, decoding falls back to regular speculative decoding.

PROPS: Progressively Private Self-alignment of Large Language Models

Differential privacy · Preference privacy · Self-alignment

Comparison of randomized-response alignment, DP-SGD alignment, and the progressive two-stage PROPS procedure.

Human preference labels can reveal sensitive information about the people providing feedback. PROPS protects these labels through a multi-stage alignment procedure in which previously privatized models help label data for subsequent stages. The framework provides preference-level privacy guarantees and improves alignment utility over DP-SGD and randomized-response baselines at matched privacy budgets.

Conformal Sparsification for Bandwidth-Efficient Edge-Cloud Speculative Decoding

Efficient inference · Conformal prediction · Edge-cloud systems

Draft token distributions are sparsified and quantized at the edge before cloud verification; feedback updates the conformal threshold.

In edge-cloud speculative decoding, sending draft-token distributions can become a communication bottleneck. This work combines sparse token distributions with lattice-based quantization, using online conformal prediction to adapt the retained token set. The analysis separates rejection caused by model mismatch from quantization distortion, and experiments evaluate the resulting latency and rejection-rate trade-offs.

CURATE: Scaling-up Differentially Private Causal Graph Discovery

Differential privacy · Causal discovery · Adaptive privacy budgets

CURATE allocates privacy budgets across successive orders of conditional-independence tests and composes their privacy guarantees.

Causal discovery makes a sequence of decisions, and privacy noise at different stages can affect the recovered graph differently. CURATE adapts the privacy budget across conditional-independence tests or score-based optimization steps while bounding cumulative privacy leakage. This allocation improves the privacy-utility trade-off over uniform budgeting in the evaluated causal-discovery settings.

Learning to Diagnose Privately: DP-Powered LLMs for Radiology Report Classification

Differential privacy · Parameter-efficient fine-tuning · Medical NLP

Public pretraining is followed by private fine-tuning of selected model parameters using locally available radiology data.

Fine-tuning language models on radiology reports can expose sensitive patient information. This work combines differentially private training with Low-Rank Adaptation to classify multiple abnormalities from report text while updating a limited set of model parameters. Experiments on MIMIC-CXR and CT-RATE examine how classification performance changes with the privacy budget.

STAMP: Selective Task-Aware Mechanism for Text Privacy

Text privacy · Task-aware privatization · Embedding perturbation

STAMP groups tokens by task relevance and privacy sensitivity, assigns group-wise budgets, and applies polar perturbation.

Text tokens differ in both privacy sensitivity and relevance to a downstream task. STAMP uses these two signals to group tokens and allocate privacy budgets, then applies a polar mechanism that perturbs embedding directions while preserving their magnitudes. Evaluations on question answering and text classification demonstrate improved privacy-utility trade-offs over the compared mechanisms.

See Publications for the full publication list.