PAPER PLAINE

Fresh research, simply explained. Updates twice daily.

Multi-type branching inference on contact trees with application to COVID-19

Mapping how diseases spread through real contact networks, not just genetic sequences.

Researchers developed a mathematical method to extract disease transmission patterns directly from contact-tracing data—who infected whom—without needing genetic sequences. The approach accounts for a key reality that older models miss: some infected people have many contacts while others have few, and this affects how fast disease spreads. When tested on COVID-19 data from India, the method accurately recovered transmission rates and contact patterns.

Public health officials use contact tracing to understand outbreak dynamics, but existing tools struggle to extract transmission rates from incomplete records. This framework turns messy contact-tracing data into precise estimates of who is most likely to spread disease and how many contacts matter, enabling faster identification of superspreaders and better targeting of interventions during future outbreaks.

Multi-Task Bayesian In-Context Learning

Teaching AI to make fast, smart predictions that adapt to new situations

Researchers developed a method that lets artificial intelligence systems quickly learn how to make predictions with built-in uncertainty estimates, even when the rules change. The approach uses a transformer model trained to read past examples and adjust its predictions for new scenarios—and it works orders of magnitude faster than traditional mathematical methods while matching their accuracy.

Machine learning systems often need to adapt predictions when conditions shift—weather forecasting when climate patterns change, medical diagnosis when treating a new population, or recommendation systems facing new user preferences. This method makes that adaptation fast enough to happen in real time while maintaining the statistical rigor that matters for high-stakes decisions. The authors demonstrated it on temperature prediction and showed it handles situations that would break less flexible approaches.

Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at Random

Evaluating AI decisions when reward data goes missing in unpredictable patterns

When hospitals or companies use past data to test new decision-making strategies, they often have incomplete records—some rewards are never recorded, others are hidden above a threshold. This creates a blind spot that breaks standard evaluation methods. The researchers developed a new statistical approach that recovers the missing information using future outcomes as clues, allowing them to fairly test new policies even when data is riddled with these gaps.

Healthcare systems and marketing platforms constantly evaluate whether new treatment or customer strategies would work better than current ones, but incomplete record-keeping undermines these tests. This method makes it possible to learn from flawed historical data without bias, meaning hospitals could confidently test new care protocols and companies could validate strategy changes using the messy real-world data they actually have.

Agentic AutoResearch forSpace Autonomy: An Auditable, LLM-Driven Research Agent for Aerospace Control Problems

AI that can teach spacecraft to fly themselves—and prove the results are real

Researchers built an AI agent that automatically designs control policies for spacecraft by proposing and testing tweaks to training code, then checking whether improvements are genuine or just statistical noise. On two docking and rendezvous problems, the AI-designed policies outperformed random parameter searches so decisively that on one task, undirected search produced no working solution at all while the AI approach succeeded every time.

Spacecraft currently rely on hand-coded control systems or policies developed through labor-intensive manual research. This framework could compress that development cycle while building in built-in verification that results are trustworthy—crucial for safety-critical aerospace applications where false confidence in a control system could end in collision or mission failure.

The Token Is a Group Element: On Lie-Algebra Attention over Matrix Lie Groups

Teaching AI to pay attention using pure geometry instead of learned rules

A new attention mechanism for AI treats tokens as geometric transformations—rotations, reflections, shearing—rather than vectors with learned features. The system scores relationships using intrinsic distance between these transformations, not learned kernels, and handles complex geometric groups (like rotations in 3D space or 2D affine transformations with scaling) that existing methods cannot. In tests on sequence completion, it matched learned approaches with 50–80 times fewer parameters and broke no geometric rules, while standard vector-based attention failed by trillions of times over.

Most AI attention mechanisms are built on learned, data-dependent rules that can violate the geometric structure they're meant to preserve. This construction builds attention directly from mathematical geometry, guaranteeing that transformations remain valid by design rather than by luck. That matters for any system working with structured spatial data—robotics, 3D vision, medical imaging, physical simulations—where breaking geometric consistency causes failures downstream.

Topological Codes Based on Space Groups

Building quantum error-correction codes with less repetitive structure

Researchers expanded how to build topological codes—a leading approach to protecting quantum computers from errors—by relaxing the requirement that they repeat perfectly across space. The new codes combine translation symmetry with rotations and reflections, and surprisingly, they can require fewer qubits in practice than the standard designs, making them simpler to build.

Quantum computers remain fragile, and error correction is essential before they can solve real problems. This work expands the toolkit for designing error-correcting codes that fit better with actual quantum hardware, potentially reducing the number of physical qubits needed to run a reliable quantum computer.

Forecasting AI-Era Productivity: The Intellectually Converged Human Framework and a Missing Cognitive Mediator in Production Function Theory

Why AI investments fail without developing workers' ability to use them

Massive spending on artificial intelligence hasn't delivered expected productivity gains because companies deploy AI without first building workers' capacity to actually use it effectively. A new framework shows that the match between AI availability and what researchers call "convergence capacity"—a combination of practical understanding, self-awareness, flexible thinking, and ability to connect ideas—accounts for 86% of productivity differences across wealthy nations, compared to just 31% for AI deployment alone.

Countries and companies are pouring billions into AI tools that sit underutilized because workers lack the cognitive skills to integrate them into their jobs. South Korea exemplifies the problem: despite strong workforce education and significant AI investment, low convergence capacity means minimal actual productivity gain. The framework suggests that before buying more AI, organizations need to invest in training that builds workers' ability to learn across domains, think flexibly, and adapt—a shift that could unlock trillions in stranded AI value currently going unrealized.

Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems

How flawed AI judges infect each other's decisions in multi-agent systems

When AI language models evaluate each other's work in team settings, their biases spread from one agent to the next—even when they're the same model. Researchers found that biased evaluators cause contagion coefficients between 0.157 and 0.352, but adding just two more evaluators to the review process cuts this bias spread by 72%, offering a simple fix.

AI systems increasingly rely on other AIs to check their work. If one model's judgment bias infects the rest of the team, bad decisions compound across the entire network. This research shows you can dramatically reduce that contamination by using evaluation committees instead of single judges—a practical safeguard for any system where AI agents depend on each other's feedback.

Pixel-Level Residual Diffusion Transformer: Scalable 3D CT Volume Generation

A faster way to generate realistic 3D medical scans from scratch

Researchers built a new AI system that can create high-resolution 3D CT scans of the chest and lungs with fine detail intact, without the computational bottlenecks that slow down existing methods. The system works in two stages: first handling large-scale structures, then filling in subtle details—an approach that outperformed competing methods on standard medical imaging benchmarks.

CT scans are expensive and expose patients to radiation, so generating realistic synthetic ones could reduce both costs and unnecessary imaging in research and clinical training. A faster, more efficient generation method means hospitals could use synthetic scans to train AI diagnostic tools and practice rare cases without scanning additional patients. This could accelerate the development of more reliable medical AI while protecting patient privacy.

AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, Knowledge Graphs, and Large Language Models

Teaching AI to explain economics using real data and tested theories

Researchers built an AI economist that generates economic reports and analyses by anchoring its claims to actual data and economic theory, rather than just producing plausible-sounding narratives. When tested on inflation forecasts and bank stress scenarios, the system produced more coherent and traceable explanations than language models working alone.

Economic analysis shapes real decisions—from Federal Reserve policy to bank lending rules—so explanations need to be trustworthy and defensible, not just fluent. This framework makes AI-generated economic reasoning transparent and checkable against actual models and evidence, reducing the risk of confident-sounding but unfounded claims influencing financial decisions.

StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs

A handful of fashion and appearance cues drive how AI judges people

AI image models make sweeping social judgments about people based on surprisingly few visual signals—mainly clothing style, age, and body type. Researchers tested six major AI systems on 25,000 carefully controlled images where only one attribute changed at a time, finding that just 15 visual cues account for nearly 80% of all the biased judgments these models make.

These AI models are already screening job applicants, assessing loan eligibility, and making other high-stakes decisions about real people. If a model judges someone's trustworthiness or earning potential based primarily on their clothes or perceived age, it can systematize discrimination at scale. This benchmark gives developers a concrete way to test and fix these specific weak points before deploying systems in consequential settings.

MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization

AI that reasons through a patient's complete medical history to guide treatment decisions

Most medical AI answers isolated questions quickly but struggles when the real answer requires connecting facts scattered across patient records, images, and sensor data. MedRLM instead builds a dynamic "evidence map" that recursively searches through a patient's full medical picture—text notes, imaging, heart rhythms, blood pressure trends, and clinical guidelines—activating deeper analysis when abnormal patterns appear, then flags cases for human review when confidence is low.

Healthcare providers in rural or under-resourced areas often lack specialists to review complex cases. A system that can systematically extract and connect evidence across all available patient data, then decide whether a case needs referral to a tertiary hospital, could reduce delays in care and improve triage accuracy. The framework's built-in uncertainty checking also prevents overconfident recommendations that might lead clinicians astray.