Adaptive Multi-Value Control in LLMs via Causal Activation Steering
LLM alignment · Activation steering · Multi-value control
LLMs often need to respond to several human values at once, and fixed steering strengths cannot adapt as a response develops. AIMES uses intermediate-layer readouts to monitor value-specific states and adjusts multiple activation interventions at each decoding step. Across model families and value combinations, it shows depth-dependent advantages over fixed joint steering while using smaller activation interventions and maintaining comparable response quality.
MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
Reward modeling · Preference learning · Data augmentation
Reward models must often learn from limited human preference data. MARS allocates more augmentation to preference pairs with small reward margins and uses semantic distance to identify pairs that need refinement before paraphrasing. Evaluations across three datasets and two reward-model backbones show improvements in average RewardBench performance and downstream alignment win rates over the evaluated augmentation baselines.
Accepted for oral presentation at Asilomar 2026. Full paper forthcoming.
Speculative speculative decoding precomputes continuations that can be reused after target-model verification. CSSD uses context-dependent rejection modeling to estimate plausible accepted lengths and online conformal calibration to select which lengths to retain in the cache. Cached continuations are used when verification produces a matching outcome; otherwise, decoding falls back to regular speculative decoding.
Human preference labels can reveal sensitive information about the people providing feedback. PROPS protects these labels through a multi-stage alignment procedure in which previously privatized models help label data for subsequent stages. The framework provides preference-level privacy guarantees and improves alignment utility over DP-SGD and randomized-response baselines at matched privacy budgets.
Conformal Sparsification for Bandwidth-Efficient Edge-Cloud Speculative Decoding
Efficient inference · Conformal prediction · Edge-cloud systems
In edge-cloud speculative decoding, sending draft-token distributions can become a communication bottleneck. This work combines sparse token distributions with lattice-based quantization, using online conformal prediction to adapt the retained token set. The analysis separates rejection caused by model mismatch from quantization distortion, and experiments evaluate the resulting latency and rejection-rate trade-offs.
Causal discovery makes a sequence of decisions, and privacy noise at different stages can affect the recovered graph differently. CURATE adapts the privacy budget across conditional-independence tests or score-based optimization steps while bounding cumulative privacy leakage. This allocation improves the privacy-utility trade-off over uniform budgeting in the evaluated causal-discovery settings.
Learning to Diagnose Privately: DP-Powered LLMs for Radiology Report Classification
Differential privacy · Parameter-efficient fine-tuning · Medical NLP
Fine-tuning language models on radiology reports can expose sensitive patient information. This work combines differentially private training with Low-Rank Adaptation to classify multiple abnormalities from report text while updating a limited set of model parameters. Experiments on MIMIC-CXR and CT-RATE examine how classification performance changes with the privacy budget.
STAMP: Selective Task-Aware Mechanism for Text Privacy
Text privacy · Task-aware privatization · Embedding perturbation
Text tokens differ in both privacy sensitivity and relevance to a downstream task. STAMP uses these two signals to group tokens and allocate privacy budgets, then applies a polar mechanism that perturbs embedding directions while preserving their magnitudes. Evaluations on question answering and text classification demonstrate improved privacy-utility trade-offs over the compared mechanisms.