Paper Digest: AISTATS 2026 Papers & Highlights
To search for papers presented at AISTATS-2026 on a specific topic, please make use of the search by venue (AISTATS-2026) service. To summarize the latest research published at AISTATS-2026 on a specific topic, you can utilize the review by venue (AISTATS-2026) service. If you are interested in browsing papers by author, we have a comprehensive list of ~ 2,000 authors (AISTATS-2026). Additionally, you may want to explore our “Best Paper” Digest (AISTATS), which lists the most influential AISTATS papers in recent years.
Since 2018, Paper Digest has built a foundation of data spanning decades of conferences, journals, and research topics. The platform features a daily digest service that sifts through tens of thousands of new papers, clinical trials, news articles, and community posts, filtering the noise to highlight what matters most to specific interests. Beyond daily updates, dozens of built-in research tools streamline the academic workflow, supporting efficient reading and writing, comprehensive literature reviews, and automated research report generation.
Paper Digest Team
New York City, New York, 10017
team@paperdigest.org
TABLE 1: Paper Digest: AISTATS 2026 Papers & Highlights
| Paper | Author(s) | |
|---|---|---|
| 1 | Enhancing LLM Safety Through A Theoretical Minimax Game Lens Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Even within English datasets, safe yet sensitive corner-case content is scarce, leading to shortcut learning by models and non-trivial false-positive rates. To mitigate these issues, we introduce a novel minimax reinforcement learning (RL) framework wherein a data generator and a classifier model co-evolve, facilitating the production of high-quality synthetic multilingual safety data. |
Yihe Deng; Yu Yang; Junkai Zhang; Wei Wang; Bo Li; |
| 2 | RL-finetuning LLMs from On- and Off-policy Data with A Single Algorithm Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel reinforcement learning algorithm (AGRO, for Any-Generation Reward Optimization) for finetuning Large Language Models. |
Yunhao Tang; Taco Cohen; David W. Zhang; Gabriel Synnaeve; Rémi Munos; |
| 3 | Low-Rank Bias, Weight Decay, and Model Merging in Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We explore the low-rank structure of the weight matrices in neural networks at the stationary points (limiting solutions of optimization algorithms) with $L2$ regularization (also known as weight decay). |
Ilja Kuzborskij; Yasin Abbasi-Yadkori; |
| 4 | Filter, Augment, Forecast: Online Data Selection for Robust Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Filter, Augment, Forecast (FAF): an online data curation strategy based on (1) data selection to filter out low-quality (e.g., noisy) examples and (2) augmentation of the remaining high-quality data. |
Ege Onur Taga; Halil Alperen Gozeten; Kutay Tire; Rahul Dalvi; Reinhard Heckel; Samet Oymak; |
| 5 | Local Inconsistency Resolution: The Interplay Between Attention and Control in Probabilistic Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a generic algorithm for learning and approximate inference with an intuitive epistemic interpretation: iteratively focus on a subset of the model and resolve inconsistencies using the parameters under control. |
Oliver Ethan Richardson; Mandana Samiei; Mehran Shakerinava; Joseph D Viviano; Abdessamad El Kabid; Ali Parviz; Yoshua Bengio; |
| 6 | Off-policy Distributional Q($\lambda$): Distributional RL Without Importance Sampling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce off-policy distributional Q($\lambda$), a new addition to the family of off-policy distributional evaluation algorithms. |
Yunhao Tang; Mark Rowland; Rémi Munos; Bernardo Avila Pires; Will Dabney; |
| 7 | Fact-Augmented Lookahead Planning for LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce LWM-Planner, a fact-augmented lookahead planning framework that improves agent behavior purely through in-context learning. |
Samuel Holt; Max Ruiz Luyten; Thomas Pouplin; Mihaela van der Schaar; |
| 8 | Demystifying Transition Matching: When and Why It Can Beat Flow Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work answers the question of when and why TM outperforms FM. |
Jaihoon Kim; Rajarshi Saha; Youngsuk Park; Minhyuk Sung; |
| 9 | Time Series Forecasting with Hahn Kolmogorov-Arnold Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose HaKAN, a versatile model based on Kolmogorov-Arnold Networks (KANs), leveraging Hahn polynomial-based learnable activation functions and providing a lightweight and interpretable alternative for multivariate time series forecasting. |
Md Zahidul Hasan; Abdessamad Ben Hamza; Nizar Bouguila; |
| 10 | Noise-Free Dynamic Rank-Adaptation Via Riemannian Methods in Federated Fine-Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We presents Riemannian LoRA algorithm with adaptive rank for federated fine-tuning of foundation models (FFT-FM), RAFFT, which resolves both issues and significantly improves the computational cost. |
Zihan Zhou; Yang Zhou; Tianshi Che; Zeru Zhang; Jiaxiang Ren; Da Yan; Zhe Jiang; yelong shen; Ruoming Jin; Jianfeng Gao; |
| 11 | Influence Attributions Can Be Systematically Altered By Model Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Influence Functions are a standard tool for attributing predictions to training data in a principled manner and are widely used in applications such as data valuation and fairness. In this work, we present realistic incentives to manipulate influence-based attributions and investigate whether these attributions can be \textit{systematically} altered by an adversary. |
Chhavi Yadav; Ruihan Wu; Kamalika Chaudhuri; |
| 12 | Hybrid Meta-Learners for Estimating Heterogeneous Treatment Effects Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce the Hybrid Learner (H-learner), a novel regularization strategy that interpolates between the direct and indirect regularizations depending on the dataset at hand. |
Zhongyuan Liang; Lars van der Laan; Ahmed Alaa; |
| 13 | The Reasoning-Creativity Trade-off: Toward Creativity-Driven Problem Solving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This can improve accuracy while still collapsing the distribution inside the correct set onto a narrow family of redundant strategies, reducing creative problem-solving. To diagnose this failure mode, we introduce Distributional Creative Reasoning (DCR), a variational framework that casts training as gradient flow on the simplex of reasoning traces. |
Max Ruiz Luyten; Mihaela van der Schaar; |
| 14 | Denoising Score Matching with Random Features: Insights on Diffusion Models From Precise Learning Curves Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We theoretically investigate the phenomena of generalization and memorization in diffusion models. |
Anand Jerry George; Rodrigo Veiga; Nicolas Macris; |
| 15 | Gaussian Approximation and Multiplier Bootstrap for Stochastic Gradient Descent Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we establish the non-asymptotic validity of the multiplier bootstrap procedure for constructing the confidence sets using the Stochastic Gradient Descent (SGD) algorithm. |
Marina Sheshukova; Sergey Samsonov; Denis Belomestny; Eric Moulines; Qi-Man Shao; Zhuo-Song Zhang; Alexey Naumov; |
| 16 | CTRLS: Chain-of-Thought Reasoning Via Latent State Transition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce CTRLS, a framework that formulates CoT reasoning as a Markov decision process (MDP) with latent state transitions, enabling explainable and state-aware exploration via distributional reinforcement learning. |
Junda Wu; Yuxin Xiong; Xintong Li; Sheldon Yu; Zhengmian Hu; Tong Yu; Rui Wang; Xiang Chen; Jingbo Shang; Julian McAuley; |
| 17 | Interpretable DNA Sequence Classification Via Dynamic Feature Generation in Decision Trees Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Contrasting these black boxes, axis-aligned decision trees offer a promising direction for interpretable DNA sequence analysis, yet they suffer from a fundamental limitation: considering individual raw features in isolation at each split limits their expressivity, which results in prohibitive tree depths that hinder both interpretability and generalization performance. We address this challenge by introducing DEFT, a novel framework that adaptively generates high-level sequence features during tree construction. |
Nicolas Huynh; Krzysztof Kacprzyk; Ryan M Sheridan; David L. Bentley; Mihaela van der Schaar; |
| 18 | Adaptive Coverage Policies in Conformal Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we leverage recent advances in e-values and post-hoc conformal inference, which allow the use of data-dependent coverage levels while maintaining valid statistical guarantees. |
Etienne Gauthier; Francis Bach; Michael I. Jordan; |
| 19 | Recency Biased Causal Attention for Time-series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a simple mechanism to introduce recency bias by reweighting attention scores with a smooth heavy-tailed decay. |
Kareem Hegazy; Michael W. Mahoney; N. Benjamin Erichson; |
| 20 | Retrieval Augmented Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: In this paper, we advocate that the dynamic and event-driven nature of time-series data makes RAG a crucial component of TSFMs and introduce a principled RAG framework for time-series forecasting, called Retrieval Augmented Forecasting (RAF). |
Kutay Tire; Ege Onur Taga; Muhammed Emrullah Ildiz; Samet Oymak; |
| 21 | Rethinking Cross-Modal Fine-Tuning: Optimizing The Interaction Between Feature Alignment and Target Fitting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing work however lacks a theoretical understanding of this critical interaction between feature alignment and target fitting. To bridge this gap, we develop a principled framework that establishes a provable generalization bound on the target error, which explains the interaction between feature alignment and target fitting through a novel concept of feature-label distortion. |
T. Khiem Tran; Manh Cuong Dao; Phi Le Nguyen; Thao Nguyen Truong; Trong Nghia Hoang; |
| 22 | PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we investigate the potential of integrating explicit periodicity modeling into the self-attention mechanism to enhance the performance of Transformer based architectures for LTSF. |
Tian Sun; Yuqi Chen; Weiwei Sun; |
| 23 | In-Context Function Learning in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Across model sizes, we find that LLM learning curves are strongly influenced by the function-generating kernels and approach the GP lower bound as the number of demonstrations increases. |
Elif Akata; Konstantinos Voudouris; Vincent Fortuin; Eric Schulz; |
| 24 | Panprediction: Optimal Predictions for Any Downstream Task and Loss Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, an emerging paradigm instead views model training as extracting enough information from data so that the model can be used to minimize many losses on many downstream tasks. We formalize a mathematical framework for this paradigm, which we call panprediction, and study its statistical complexity. |
Sivaraman Balakrishnan; Nika Haghtalab; Daniel Hsu; Brian W Lee; Eric Zhao; |
| 25 | Multi-Armed Sampling Problem and The End of Exploration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces the framework of multi-armed sampling, which serves as the sampling counterpart to the optimization problem of multi-armed bandits. |
Mohammad Pedramfar; Siamak Ravanbakhsh; |
| 26 | MineGrad: Gradient Inversion Attacks on LoRA Fine-Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we investigate gradient inversion attacks on LoRA fine-tuning. |
Hasin Us Sami; Swapneel Sen; Basak Guler; |
| 27 | Near-Optimal Clustering in Mixture of Markov Chains Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study the problem of clustering $T$ trajectories of length $H$, each generated by one of $K$ unknown ergodic Markov chains over a finite state space of size $S$. |
Junghyun Lee; Yassir Jedra; Alexandre Proutiere; Se-Young Yun; |
| 28 | GL-LowPopArt: A Nearly Instance-Wise Minimax-Optimal Estimator for Generalized Low-Rank Trace Regression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present **GL-LowPopArt**, a novel Catoni-style estimator for generalized low-rank trace regression.We also introduce a new problem setting called **bilinear dueling bandits**, a contextualized version of dueling bandits with a general preference model. |
Junghyun Lee; Kyoungseok Jang; Kwang-Sung Jun; Milan Vojnovic; Se-Young Yun; |
| 29 | Shift Is Good: Mismatched Data Mixing Improves Test Performance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We consider training and testing on mixture distributions with different training and test proportions. |
Marko Medvedev; Kaifeng Lyu; Zhiyuan Li; Nathan Srebro; |
| 30 | AMRM-Pure: Semantic-Preserving Adversarial Purification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our theoretical and experimental analysis reveals that AMRM is highly sensitive to adversarial noise, as such noise significantly distorts patch relationships. Based on this observation, we propose AMRM-Pure, a purification framework that denoises adversarial inputs by preserving patch-level semantics, and formulate this process as a tractable optimization problem with respect to the input. |
Zhihao Dou; Zhiqiang Gao; Dongfei Cui; Weida Wang; Qinjian Zhao; Dinggen Zhang; Jun Yan; Zeke Xie; Shufei Zhang; |
| 31 | On The Calibration of Survival Models with Competing Risks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that existing calibration measures are not suited to the competing-risk setting and that recent models do not give well-behaved probabilities. To address this, we introduce a dedicated framework with two novel calibration measures that are minimized for oracle estimators (*i.e.*, both measures are proper). |
Julie Alberge; Tristan Haugomat; Gaël Varoquaux; Judith Abécassis; |
| 32 | Neural Additive Experts: Context-Gated Experts for Controllable Model Additivity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Introducing feature interactions can boost accuracy yet may obscure individual feature contributions. To address these issues, we propose Neural Additive Experts (NAEs), a novel framework that seamlessly balances interpretability and accuracy. |
Guangzhi Xiong; Sanchit Sinha; Aidong Zhang; |
| 33 | Auto-Regressive Masked Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Masked diffusion models (MDMs) have emerged as a promising approach for language modeling, yet they face a performance gap compared to autoregressive models (ARMs) and require more training iterations. In this work, we present the Auto-Regressive Masked Diffusion (ARMD) model, an architecture designed to bridge this gap by unifying the training efficiency of autoregressive models with the strengths of diffusion-based learning. |
Mahdi Karami; Ali Ghodsi; |
| 34 | Learning to Choose or Choosing to Learn: Best-of-N Vs. Supervised Fine-Tuning for Bit String Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Using the bit string generation problem as a case study, we theoretically compare two standard methods for adapting large language models to new tasks. |
Seamus Somerstep; Vinod Raman; Unique Subedi; Yuekai Sun; |
| 35 | Understanding Generalization in Node and Link Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce a unified framework for analyzing the generalization properties of MPNNs in inductive and transductive node and link prediction settings, incorporating diverse architectural parameters and loss functions, and quantifying the influence of graph structure. |
Antonis Vasileiou; Timo Stoll; Christopher Morris; |
| 36 | Structured Temporal Inference in State-Space Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a framework for structured temporal inference in nonlinear state-space models (SSMs) with hybrid latent dynamics that mix discrete and continuous variables. |
Hamidreza Hashempoor; |
| 37 | Doctor Rashomon and The UNIVERSE of Madness: Variable Importance with Unobserved Confounding and The Rashomon Effect Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Variable importance (VI) methods are often used for hypothesis generation, feature selection, and scientific validation. |
Jon Donnelly; Srikar Katta; Emanuele Borgonovo; Cynthia Rudin; |
| 38 | Aggregation on Learnable Manifolds for Asynchronous Federated Optimisation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Asynchronous federated learning (FL) with heterogeneous clients faces two key issues: curvature-induced loss barriers encountered by standard linear parameter interpolation techniques (e.g. FedAvg) and interference from stale updates misaligned with the server’s current optimisation state. To alleviate these issues, we introduce a geometric framework that casts aggregation as curve learning in a Riemannian model space and decouples choice of update direction from staleness conflict resolution. |
Archie Licudi; Anshul Thakur; Soheila Molaei; Danielle Belgrave; David A. Clifton; |
| 39 | A Geometric Approach to Optimal Experimental Design Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel geometric framework for optimal experimental design (OED). |
Gavin Kerrigan; Christian A. Naesseth; Tom Rainforth; |
| 40 | A Proof of Learning Rate Transfer Under $\mu$P Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We provide the first proof of learning rate transfer with width in a linear multi-layer perceptron (MLP) parametrized with $\mu$P, a neural network parameterization designed to “maximize” feature learning in the infinite-width limit. |
Soufiane Hayou; |
| 41 | Conformal Prediction in Hierarchical Classification with Constrained Representation Complexity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we extend the split conformal prediction framework to hierarchical classification, where prediction sets are commonly restricted to internal nodes of a predefined hierarchy, and propose two computationally efficient inference algorithms. |
Thomas Mortier; Alireza Javanmardi; Yusuf Sale; Eyke Hüllermeier; Willem Waegeman; |
| 42 | Best Policy Learning From Trajectory Preference Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose Posterior Sampling for Preference Learning ($\mathsf{PSPL}$), a novel algorithm inspired by Top-Two Thompson Sampling that maintains posteriors over the reward model and dynamics. |
Akhil Agnihotri; Rahul Jain; Deepak Ramachandran; Zheng Wen; |
| 43 | Bad Values But Good Behavior: Learning Highly Misspecified Bandits with Function Approximation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Function approximation with parametric, feature-based reward models is widely used to enable decision-making in bandits with large action spaces. |
Debangshu Banerjee; Aditya Gopalan; |
| 44 | A Scalable Lift-and-Project Differentiable Approach For The Maximum Cut Problem Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a scalable framework for solving the Maximum Cut (MaxCut) problem in large graphs using projected gradient ascent on quadratic objectives. |
Ismail Alkhouri; Mian Wu; CUNXI YU; Jia Liu; Rongrong Wang; Alvaro Velasquez; |
| 45 | Neural Doubly Robust Proximal Causal Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a new theoretical bound on the estimation accuracy of the treatment bridge, and we analyze the variance of the doubly-robust estimator. |
Ruolin Meng; Dhanajit Brahma; Ricardo Henao; Lawrence Carin; |
| 46 | Differentially Private E-Values Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To ensure their safe release, we propose a general framework for differentially private e-values that transforms any non-private e-value into a differentially private one. |
Daniel Csillag; Diego Mesquita; |
| 47 | Conservative Inference in Switchback Experiments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the standard pipeline of estimating the average treatment effect (ATE) with a difference-in-means (DM) estimator can exhibit systematic bias in dynamic settings with evolving system state, due to intertemporal dependence (“carryover effects). In this paper, we study this bias in a continuous-time Markov chain model of switchback experiments with stochastically monotone dynamics and state-monotone rewards; these are reasonable representations of mean-reverting and auto-regressive systems. |
Jose Blanchet; Peter Glynn; Ramesh Johari; Linjia Wu; Wenqian Xing; |
| 48 | Creator Incentives in Recommender Systems: A Cooperative Game-Theoretic Approach for Stable and Fair Collaboration in Multi-Agent Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For heterogeneous agents, the game still admits a non-empty core, though convexity and Shapley value core-membership are no longer guaranteed. To address this, we propose a simple regret-based payout rule that satisfies three out of the four Shapley axioms and also lies in the core. |
Ramakrishnan K; Arpit Agarwal; Lakshmi Subramanian; Maximilian Nickel; |
| 49 | Structured Matrix Scaling for Multi-Class Calibration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Post-hoc recalibration methods are widely used to ensure that classifiers provide faithful probability estimates. |
Eugène Berta; David Holzmüller; Michael I. Jordan; Francis Bach; |
| 50 | Beyond The Ideal: Analyzing The Inexact Muon Update Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop our analysis within the general framework of Linear Minimization Oracle (LMO)-based optimization, introducing a realistic additive error model to capture the inexactness of practical approximation schemes. |
Egor Shulgin; Sultan AlRashed; Peter Richtárik; Francesco Orabona; |
| 51 | Learning Equivariant Functions Via Quadratic Forms Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we introduce a method for learning group (known or unknown) equivariant functions by learning the associated quadratic form $x^T A x$ corresponding to the group from the data. |
Pavan Karjol; Vivek V Kashyap; Rohan Kashyap; Prathosh AP; |
| 52 | Beyond ReLU: How Activations Affect Neural Kernels and Random Wide Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the property of these kernels are poorly understood for activation functions other than powers of the ReLU. Our main contribution is a characterization of the RKHS of these kernels for activation functions whose only non-smoothness is at zero. |
David Holzmüller; Max Schölpple; |
| 53 | Hyperbolic Learning with Supervision from Any Granularity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a coarse-to-fine Busemann approach, where images are optimized to the correct region of the hyperbolic embedding space by projecting their labels — which can be as precise or generic as desired — to ideal prototypes on the boundary of the Poincaré ball. |
Mina Ghadimi Atigh; Max van Spengler; Teng Long; Melika Ayoughi; Tejaswi Kasarla; Pascal Mettes; |
| 54 | The Majority Vote Paradigm Shift: When Popular Meets Optimal Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, despite its importance, the optimality of MV’s label aggregation has not been extensively studied. We address this gap in our work by characterising the conditions under which MV achieves the theoretically optimal lower bound on label estimation error. |
Antonio Purificato; Maria Sofia Bucarelli; Anil Kumar Nelakanti; Andrea Bacciu; Fabrizio Silvestri; Amin Mantrach; |
| 55 | Weighted Quantization Using MMD: From Mean Field to Mean Shift Using Gradient Flows Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Formally, we seek a weighted mixture of Dirac measures that best approximates the target distribution. |
Ayoub Belhadji; Daniel Sharp; Youssef Marzouk; |
| 56 | Conditional Vendi Score: Prompt-Aware Diversity Evaluation for Generative AI Models and LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For Conditional-Vendi, we introduce a truncated-spectrum approximation that yields scalable and consistent estimates. |
Mohammad Jalali; Azim Ospanov; Amin Gohari; Farzan Farnia; |
| 57 | Semi-Random Noisy and One-Bit Matrix Completion Via Nonconvex Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we give a unified framework for semi-random matrix recovery applicable to a broad family of observation models. |
Xing Gao; Binhao Chen; Yu Cheng; |
| 58 | Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The average reward is a fundamental performance metric in reinforcement learning (RL) focusing on the long-run performance of an agent. Differential temporal difference (TD) learning algorithms are a major advance for average reward RL as they provide an efficient online method to learn the value functions associated with the average reward in both on-policy and off-policy settings. |
Ethan Blaser; Jiuqi Wang; Shangtong Zhang; |
| 59 | On The Interplay of Priors and Overparametrization in Bayesian Neural Network Posteriors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study how overparametrization and priors together reshape BNN posteriors and derive implications allowing us to better understand their interplay. |
Julius Kobialka; Emanuel Sommer; Chris Kolb; Juntae Kwon; Daniel Dold; David Rügamer; |
| 60 | Spectral Thresholds in Correlated Spiked Models and Fundamental Limits of Partial Least Squares Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We provide a rigorous random matrix theory analysis of spiked cross-covariance models where the signals across two high-dimensional data channels are partially aligned. |
Pierre Mergny; Lenka Zdeborová; |
| 61 | VIPaint: Image Inpainting with Pre-Trained Diffusion Models Via Variational Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a hierarchical variational inference algorithm that optimizes a non-Gaussian Markov approximation of the true diffusion posterior. |
Sakshi Agarwal; Gabriel Hope; Jimin Heo; Erik B. Sudderth; |
| 62 | Loss-Driven Bayesian Active Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a rigorous loss-driven approach to Bayesian active learning that allows data acquisition to directly target the loss associated with a given decision problem. |
Zhuoyue Huang; Freddie Bickford Smith; Tom Rainforth; |
| 63 | Gradient Descent with Provably Tuned Learning-rate Schedules Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we develop novel analytical tools for provably tuning hyperparameters in gradient-based algorithms that apply to non-convex and non-smooth functions. |
Dravyansh Sharma; |
| 64 | Corruption-robust Offline Multi-agent Reinforcement Learning from Human Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We consider robustness against data corruption in offline multi-agent reinforcement learning from human feedback (MARLHF) under a strong‐contamination model: given a dataset $D$ of trajectory–preference tuples (each preference being an $n$-dimensional binary label vector representing each of the $n$ agents’ preferences), an $\epsilon$-fraction of the samples may be arbitrarily corrupted. |
Andi Nika; Debmalya Mandal; Parameswaran Kamalaruban; Adish Singla; Goran Radanovic; |
| 65 | SenTSR-Bench: Thinking with Injected Knowledge for Time-Series Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Conversely, fine-tuned time-series LLMs (TSLMs) understand these patterns but lack the capacity to generalize reasoning for more complicated questions. To bridge this gap, we propose a hybrid knowledge-injection framework that injects TSLM-generated insights directly into GRLM’s reasoning trace, thereby achieving strong time-series reasoning with in-domain knowledge. |
Zelin He; Boran Han; Xiyuan Zhang; Shuai Zhang; Haotian Lin; Qi Zhu; Haoyang Fang; Danielle C. Maddix; Abdul Fatir Ansari; Akash Chandrayan; Abhinav Pradhan; Bernie Wang; Matthew Reimherr; |
| 66 | Modeling Multi-Objective Tradeoffs with Monotonic Utility Functions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel, principled two-step process for obtaining a compact set of PO points that aligns with user preferences, which are specified a priori as general monotonic utility functions (MFs). |
Edward Chen; Natalie Dullerud; Thomas Niedermayr; Elizabeth Kidd; Ransalu Senanayake; Pang Wei Koh; Sanmi Koyejo; Carlos Guestrin; |
| 67 | Calibrated Predictive Lower Bounds on Time-to-Unsafe-Sampling in LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce time-to-unsafe-sampling, a novel safety measure for generative models, defined as the number of generations required by a large language model (LLM) to trigger an unsafe (e.g., toxic) response. |
Hen Davidov; Shai Feldman; Gilad Freidkin; Yaniv Romano; |
| 68 | Adversarial Robustness in One-Stage Learning-to-Defer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the first framework for adversarial robustness in one-stage L2D, covering both classification and regression. |
Yannis Montreuil; Yu Letian; Axel Carlier; Lai Xing Ng; Wei Tsang Ooi; |
| 69 | Online Learning-to-Defer with Varying Experts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the first online L2D algorithm for multiclass classification with bandit feedback and a dynamically varying pool of experts. |
Yannis Montreuil; Hoang Duy Dang; Maxime Meyer; Lai Xing Ng; Axel Carlier; Wei Tsang Ooi; |
| 70 | Optimal Query Allocation in Extractive QA with LLMs: A Learning-to-Defer Framework with Theoretical Guarantees Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a Learning-to-Defer framework that routes EQA queries across a pool of models with varying capabilities and costs to balance accuracy and efficiency. |
Yannis Montreuil; Yeo Shu Heng; Axel Carlier; Lai Xing Ng; Wei Tsang Ooi; |
| 71 | Amortized In-Context Mixed Effect Transformer Models: A Zero-Shot Approach for Pharmacokinetics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present the Amortized In-Context Mixed-Effect Transformer (AICMET) model, a transformer‑based, latent‑variable framework that unifies mechanistic compartmental priors with amortized, in‑context Bayesian inference. |
Cesar Ojeda; Ramses J Sanchez; Wilhelm Huisinga; Niklas Hartung; |
| 72 | LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present LLMPhy, a black-box optimization framework that integrates large language models (LLMs) with physics simulators for physical reasoning.As existing physical reasoning benchmarks rarely account for parameter identifiability, we introduce three new datasets designed to evaluate physical reasoning in zero-shot settings. |
Anoop Cherian; Radu Corcodel; Siddarth Jain; Diego Romeres; |
| 73 | The Cross-Context Threshold Test: Detecting Discrimination Under Environmental Shifts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a cross-context threshold test that enforces distributional invariance and monotonic threshold decay. |
Jun Yuan; Xinyue Ye; |
| 74 | FastRank: Fast Tensor Rank Approximation Based on Spectral Energy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce $FastRank$, a theoretically grounded method that estimates rank without CPD computation. |
Konstantinos Bougiatiotis; Georgios Paliouras; |
| 75 | Discrete State Diffusion Models: A Sample Complexity Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present a principled theoretical framework for discrete-state diffusion, providing the first sample complexity bound of $\widetilde{\mathcal{O}}(\epsilon^{-2})$. |
Aadithya Srikanth; Mudit Gaur; Vaneet Aggarwal; |
| 76 | Generalization Bounds Under Heavy-Tailed Losses Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we analyze the generalization error under the heavy-tailed assumption on the loss function with respect to the data-generating distribution. |
Gholamali Aminian; |
| 77 | Zeroth-Order Stochastic Compositional Gradient Descent: Towards Black-Box Sparse AUC Maximization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A central challenge arises from integrating zeroth-order gradient estimation with hard-thresholding operators in the compositional framework, which has remained unresolved. To overcome this difficulty, we propose the Zeroth-Order Stochastic Compositional Hard-Thresholding (ZO-SCHT) algorithm, which, to the best of our knowledge, is the first method for black-box sparse AUC maximization. |
Wenkang Wang; Dongxu Liu; Bin Gu; |
| 78 | Rethinking Probabilistic Circuit Parameter Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our analysis reveals a fundamental issue that existing mini-batch EM and gradient-based methods fail to properly regularize distribution changes, causing each update to effectively overfit the current mini-batch. Motivated by this insight, we introduce anemone, a new mini-batch EM algorithm for PCs. |
Anji Liu; Zilei Shao; Guy Van den Broeck; |
| 79 | Visual Prompting Reimagined: The Power of Activation Prompts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, there exists a noticeable performance gap between VP and conventional fine-tuning methods, highlighting an unexplored realm in theory and practice to understand and advance (input-level) VP to reduce its current performance gap. Towards this end, we introduce a generalized concept, termed activation prompt (AP), which extends the scope of (input-level) VP by enabling universal perturbations to be applied to activation maps within the intermediate layers of the model. |
Yihua Zhang; Hongkang Li; Yuguang Yao; Aochuan Chen; Shuai Zhang; Pin-Yu Chen; Meng Wang; Sijia Liu; |
| 80 | On Computational Limits of FlowAR Models: Expressivity and Efficiency Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This gap limits our understanding of their inherent computational limits and practical efficiency. In this study, we address this gap by analyzing the circuit complexity of the FlowAR architecture. |
Yang Cao; Chengyue Gong; Yekun Ke; Xiaoyu Li; Yingyu Liang; Zhizhou Sha; Zhenmei Shi; Zhao Song; |
| 81 | Provable Affine Identifiability of Nonlinear CCA Under Latent Distributional Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we establish the sufficient conditions under which nonlinear Canonical Correlation Analysis (CCA) recovers ground-truth latent factors up to an affine transformation. |
Zhiwei Han; Stefan Matthes; Hao Shen; |
| 82 | Tractable Uncertainty-Aware Meta-Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we present LUMA, a meta-learning method for regression that (1) makes probabilistic predictions on in-distribution tasks efficiently, (2) is capable of detecting OoD context data, and (3) handles heterogeneous, multimodal task distributions effectively. |
Young-Jin Park; Cesar Almecija; Apoorva Sharma; Navid Azizan; |
| 83 | Exact Tensor Completion Beyond Isotropy and Invertibility Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, a tensor completion problem is studied, which aims to perfectly recover the tensor from partial observations. |
Li Ge; Lin Chen; Yudong Chen; Xue Jiang; |
| 84 | PolarQuant: Vector Quantization with Polar Transformation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce PolarQuant, a novel online vector quantization method that leverages random preconditioning and polar transformation. |
Insu Han; Praneeth Kacham; Amin Karbasi; Vahab Mirrokni; Amir Zandieh; |
| 85 | Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, DEQs face the challenge of requiring vastly more computational time for training and inference than conventional methods, as they repeatedly perform fixed-point iterations with no convergence guarantee upon each input. Therefore, this study explored an approach to improve fixed-point convergence and consequently reduce computational time by restructuring the model architecture to guarantee fixed-point convergence. |
Naoki Sato; Hideaki Iiduka; |
| 86 | Representation Learning Via Non-Contrastive Mutual Information Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we aim to develop a self-supervised objective that combines the strength of both types. |
Zhaohan Daniel Guo; Bernardo Avila Pires; Khimya Khetarpal; Dale Schuurmans; Bo Dai; |
| 87 | ACE-KT: Cascaded Cognitive Modeling for Stage-wise Knowledge Tracing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, modeling student learning as a process of cognitive transformation rather than as a mere sequence of time-stamped events remains a fundamental challenge in KT research. To address this issue, we propose **ACE-KT** (c**A**scaded **C**ognitive mod**E**ling for **K**nowledge **T**racing), a novel framework inspired by cognitive process theory, which shifts the focus from purely sequential modeling to cognitive representation learning. |
Teng Guo; Yubin Xia; Jinsen Ke; Mingliang Hou; Jiaqi Zheng; Zitao Liu; |
| 88 | I-IF-Learn: Iterative Feature Selection and Unsupervised Learning for High-Dimensional Complex Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose i-IF-Learn, an iterative unsupervised framework that jointly performs feature selection and clustering. |
Chen Ma; Wanjie Wang; Shuhao Fan; |
| 89 | Deep Polynomial Chaos Expansion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The applicability of PCE to high-dimensional problems is limited by poor scalability, as the number of basis functions grows exponentially with the number of parameters. In this paper, we address this challenge by combining PCE with ideas from tractable probabilistic circuits, resulting in *deep polynomial chaos expansion* (DeepPCE)—a deep generalization of PCE that scales effectively to high-dimensional input spaces. |
Johannes Exenberger; Sascha Ranftl; Robert Peharz; |
| 90 | A Covering Framework for Offline POMDPs Learning Using Belief Space Metric Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces a novel covering analysis framework that exploits the intrinsic metric structure of the belief space (distributions over latent states) to relax traditional coverage assumptions. |
Youheng Zhu; Yiping Lu; |
| 91 | Dashed Line Defense: Plug-And-Play Defense Against Adaptive Score-Based Query Attacks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By introducing ambiguity in how the observed loss reflects the true adversarial strength of candidate examples, DLD prevents attackers from reliably analyzing and adapting their queries, effectively disrupting the AE generation process. We provide theoretical guarantees of DLD’s defense capability and validate its effectiveness through experiments on ImageNet, demonstrating that DLD consistently outperforms prior defenses—even under worst-case adaptive attacks—while preserving the model’s predicted labels. |
Yanzhang Fu; Jizhou Luo; Zizheng Guo; |
| 92 | Rethinking Intrinsic Dimension Estimation in Neural Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we highlight a crucial discrepancy between theory and practice of IDs in neural representations, theoretically and empirically showing that common ID estimators are, in fact, not tracking the true underlying ID of the representation. |
Rickmer Schulte; David Rügamer; |
| 93 | Provably Efficient and Agile Randomized Q-Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel variant of Q-learning algorithm, referred to as RandomizedQ, which integrates sampling-based exploration with agile, step-wise, policy updates, for episodic tabular RL. |
He Wang; Xingyu Xu; Yuejie Chi; |
| 94 | TabTreeFormer: Tabular Data Generation Using Hybrid Tree-Transformer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose TabTreeFormer, a hybrid transformer architecture that integrates inductive biases of tree-based models (e.g., non-smoothness and non-rotational invariance) to effectively handle the discrete and weakly correlated features in tabular datasets. |
Jiayu Li; Bingyin Zhao; Zilong Zhao; Uzair Javaid; Biplab Sikdar; |
| 95 | Likelihood-Free Inference Via Structured Score Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a likelihood-free inference framework that combines score matching with gradient-based optimization and bootstrap procedures to facilitate parameter estimation together with uncertainty quantification. |
Haoyu Jiang; Yuexi Wang; Yun Yang; |
| 96 | SFBD Flow: A Continuous-Optimization Framework for Training Diffusion Models with Noisy Samples Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We reinterpret SFBD as an alternating projection algorithm and introduce a continuous variant, SFBD flow, that removes the need for alternating steps. |
Haoye Lu; Darren Lo; Yaoliang Yu; |
| 97 | Meet Me at The Arm: The Cooperative Multi Armed Bandits Problem with Shareable Arms Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This setting generalizes the classical unit-capacity model and introduces new challenges in coordination and capacity discovery under severe feedback limitations. We propose A-CAPELLA (Algorithm for Capacity-Aware Parallel Elimination for Learning and Allocation), a decentralized learning algorithm that achieves logarithmic regret in this generalized regime via protocol-driven coordination. |
Xinyi Hu; Aldo Pacchiano; |
| 98 | Private and Efficient Federated Statistical Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a novel federated statistical learning framework that ensures efficient, robust, and privacy-preserving estimation. |
Jaemu Heo; Xiwen Feng; Jeonghun Kang; Taehwan Kim; Changgee Chang; |
| 99 | Learning Physical Operators Using Neural Operators Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Neural operators have emerged as promising surrogate models for solving partial differential equations (PDEs), but struggle to generalise beyond training distributions and are often constrained to a fixed temporal discretisation. This work introduces a physics-informed training framework that addresses these limitations by decomposing PDEs using operator splitting methods, training separate neural operators to learn individual non-linear physical operators while approximating linear operators with fixed finite-difference convolutions. |
Vignesh Gopakumar; Ander Gray; Daniel Giles; Lorenzo Zanisi; Matt J. Kusner; Timo Betcke; Stanislas Pamela; Marc Peter Deisenroth; |
| 100 | Inverse-Free Sparse Variational Gaussian Processes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We make the inverse-free approach practical by proposing a better-conditioned bound and deriving a matmul-only natural-gradient update for the auxiliary parameter, markedly improving stability and convergence. |
Stefano Cortinovis; Laurence Aitchison; Stefanos Eleftheriadis; Mark van der Wilk; |
| 101 | Near-Optimal Sample Complexities of Divergence-based S-rectangular Distributionally Robust Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study the empirical value iteration algorithm for divergence-based S-rectangular DR-RL and establish near-optimal sample complexity bounds of $\widetilde{O}(|\mathcal{S}||\mathcal{A}|(1-\gamma)^{-4}\varepsilon^{-2})$, where $\varepsilon$ is the target accuracy, $|\mathcal{S}|$ and $|\mathcal{A}|$ denote the cardinalities of the state and action spaces, and $\gamma$ is the discount factor. |
Zhenghao Li; Shengbo Wang; Nian Si; |
| 102 | High Effort, Low Gain: Fundamental Limits of Active Learning for Linear Dynamical Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we present sample complexity lower bounds that capture the choice of the selected excitation input. |
Nicolas Chatzikiriakos; Kevin Jamieson; Andrea Iannelli; |
| 103 | Incentivizing Truthful Submissions in A Data Marketplace for Mean Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We aim to maximize welfare or profit under key constraints: individual rationality for buyers and contributors, incentive compatibility (contributors are incentivized to comply with data collection instructions and truthfully report the collected data), and budget balance (total contributor payments equals total revenue). |
Keran Chen; Alex Clinton; Kirthevasan Kandasamy; |
| 104 | Quantifying Epistemic Uncertainty in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a method based on Fisher information that explicitly isolates epistemic variance, producing more reliable plausibility scores for generated data. |
Aditi Gupta; Raphael A Meyer; Yotam Yaniv; Elynn Chen; N. Benjamin Erichson; |
| 105 | Deliberate-When-Needed: Flow-Reasoner for Neuro-Symbolic Continuous Thought Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Flow-Reasoner, a Deliberate-When-Needed neuro-symbolic model that integrates continuous latent cognition with selective symbolic reasoning. |
Wenjie Shen; Boyang Li; Chao Yang; Shuang Li; |
| 106 | We Still Don’t Understand High-Dimensional Bayesian Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing high-dimensional Bayesian optimization (BO) methods aim to overcome the curse of dimensionality by carefully encoding structural assumptions, from locality to sparsity to smoothness, into the optimization procedure. |
Colin Doumont; Donney Fan; Natalie Maus; Jacob R. Gardner; Henry Moss; Geoff Pleiss; |
| 107 | LLMs Judging LLMs: A Simplex Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While this is justified if judges are perfectly accurate, it is unclear when such an approach is theoretically valid and practically robust. We study these questions for the task of ranking LLM candidates from a novel geometric perspective: for $M$-level scoring systems, both LLM judges and candidates can be represented as points on an $(M-1)$-dimensional probability simplex, where geometric concepts (e.g., triangle areas)correspond to key ranking concepts. |
Patrick Vossler; Fan Xia; Yifan Mai; Adarsh Subbaswamy; Jean Feng; |
| 108 | Impact of Positional Encoding: Clean and Adversarial Rademacher Complexity for Transformers Under In-Context Regression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we provide the first generalization analysis for single-layer Transformer under in-context regression that explicitly accounts for a trainable PE module. |
Weiyi He; Yue Xing; |
| 109 | Lloyd’s $K$-Means Clustering Algorithm Is Frank-Wolfe in Disguise Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we establish a novel connection between Lloyd’s algorithm and the Frank-Wolfe (FW) algorithm, a prominent first-order method for projection-free optimization. |
Michael Pokojovy; J. Marcus Jobe; Simon Lacoste-Julien; |
| 110 | Preconditioned Attention: Enhancing Efficiency in Transformer Blocks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This ill-conditioning is a well-known obstacle for gradient-based optimizers, leading to inefficient training. To address this issue, we introduce preconditioned attention, a novel approach that incorporates a conditioning matrix into each attention head. |
Hemanth Saratchandran; |
| 111 | Efficient Learning of Stationary Diffusions with Stein-type Discrepancies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Leveraging the connection between KDS and Stein discrepancies, we introduce the Stein-type KDS (SKDS) as an alternative formulation. |
Fabian Bleile; Sarah Lumpp; Mathias Drton; |
| 112 | Auditing Pay-Per-Token in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we develop an auditing framework based on martingale theory that enables a trusted third-party auditor who sequentially queries a provider to detect token misreporting. |
Ander Artola Velasco; Stratis Tsirtsis; Manuel Gomez Rodriguez; |
| 113 | Boosted GFlowNets: Improving Exploration Via Sequential Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Nonetheless, in practice, they often struggle to explore the reward landscape evenly: trajectories toward easy-to-reach regions dominate training, while hard-to-reach modes receive vanishing or uninformative gradients, leading to poor coverage of high-reward areas. We address this imbalance with Boosted GFlowNets, a method that sequentially trains an ensemble of GFlowNets, each optimizing a residual reward that compensates for the mass already captured by previous models. |
Pedro Dall’Antonia; Tiago Silva; Daniel Augusto de Souza; César Lincoln Mattos; Diego Mesquita; |
| 114 | Understanding The Benefits of SimCLR Pre-Training in Two-Layer Convolutional Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a theoretical case study of SimCLR. |
Han Zhang; Yuan Cao; |
| 115 | Accelerating PDE Surrogates Via RL-Guided Mesh Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Deep learning–based surrogate models for parametric partial differential equations (PDEs) can deliver high-fidelity approximations but remain prohibitively data-hungry: training often requires thousands of fine-grid simulations, each incurring substantial computational cost. To address this challenge, we introduce RLMesh, an end-to-end framework for efficient surrogate training under limited simulation budget. |
Yang Meng; Ruoxi Jiang; Zhuokai Zhao; Chong Liu; Rebecca Willett; Yuxin Chen; |
| 116 | A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce into the lower-level RL objective an attenuating entropy regularization, which enables asymptotically unbiased upper-level hyper-gradient estimation without solving the unregularized RL problem exactly. |
Sihan Zeng; Sujay Bhatt; Sumitra Ganesh; Alec Koppel; |
| 117 | Regularized Operator Extrapolation Method For Stochastic Hierarchical Variational Inequality Problems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Regularized Operator Extrapolation (R-OpEx), a single-loop first-order algorithm for smooth and nonsmooth BVIs with stochastic monotone operators. |
Mohammad Khalafi; Digvijay Boob; |
| 118 | Proof of The TAP Free Energy for High-Dimensional Linear Regression with Spherical Priors at All Temperatures Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Approximate inference is central to Bayesian learning, with variational inference (VI) providing a scalable framework for posterior approximation. |
Zhiyuan Yu; Jingbo Liu; |
| 119 | Accelerating Byzantine-Robust Distributed Learning with Compressed Communication Via Double Momentum and Variance Reduction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Byz-DM21, a novel Byzantine-robust and communication-efficient stochastic distributed learning algorithm. |
Yanghao Li; Changxin Liu; Yuhao Yi; |
| 120 | Linearly Separable Features in Shallow Nonlinear Networks: Width Scales Polynomially with Intrinsic Data Dimension Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, these findings often lack rigorous justifications, even under relatively simple settings. In this work, we address this gap by examining the linear separation capabilities of shallow nonlinear networks. |
Alec S. Xu; Can Yaras; Peng Wang; Qing Qu; |
| 121 | Semi-Implicit Variational Inference Via Kernelized Path Gradient Descent Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a kernelized KL divergence estimator that stabilizes training through nonparametric smoothing, effectively addressing the blindness” challenge. |
Tobias Pielok; Bernd Bischl; David Rügamer; |
| 122 | FIELDING: Clustered Federated Learning with Data Drift Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose FIELDING, a CFL framework for handling diverse types of data drift with low overhead. |
Minghao Li; Dmitrii Avdiukhin; Rana Shahout; Nikita Ivkin; Vladimir Braverman; Minlan Yu; |
| 123 | Do We Need Rebalancing Strategies? A Theoretical and Empirical Study Around SMOTE and Its Variants Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: In this paper, we derive several non-asymptotic upper bound on SMOTE density. |
Abdoulaye SAKHO; Emmanuel Malherbe; Erwan Scornet; |
| 124 | A Continuous Time Markov Chain Framework for Insertion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we derive a diffusion-style denoising objective for ILMs from first principles by formulating the noising process as a continuous-time Markov chain on the space of variable-length sequences. |
Dhruvesh Patel; Benjamin Rozonoyer; Soumitra Das; Tahira Naseem; Tim G. J. Rudner; Andrew McCallum; |
| 125 | How to Approximate Inference with Subtractive Mixture Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, how to effectively use SMMs for VI and IS is still an open question as they do not provide latent variable semantics and therefore cannot use sampling schemes for classical MMs. In this work, we study how to circumvent this issue by designing several expectation estimators for IS and learning schemes for VI with SMMs, and we empirically evaluate them for distribution approximation. |
Lena Zellinger; Nicola Branchini; Lennert De Smet; Víctor Elvira; Nikolay Malkin; Antonio Vergari; |
| 126 | MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: With increasing size of large language models (LLMs), full-parameter fine-tuning imposes substantial memory demands. To alleviate this, we propose a novel memory-efficient training paradigm called Momentum Low-rank compression (MLorc). |
Wei Shen; Zhang Yaxiang; Minhui Huang; Mengfan Xu; Jiawei Zhang; Cong Shen; |
| 127 | Eliciting Truthful Feedback for Preference-Based Learning Via The VCG Mechanism Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This setting poses two main challenges: (i) the agents’ cost functions may be unknown to them or difficult to specify explicitly, and (ii) agents may misreport their costs strategically. To address these challenges, we propose an algorithm that combines preference-based learning with Vickrey–Clarke–Groves (VCG) payments to incentivize truthful reporting. |
Leo Landolt; Anna Maria Maddux; Andreas Schlaginhaufen; Saurabh Vaishampayan; Maryam Kamgarpour; |
| 128 | Brenier Isotonic Regression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We consider a multi-output regression problem where a regression function is cyclically monotone. |
Han Bao; Amirreza Eshraghi; Yutong Wang; |
| 129 | Three-Step Nav: A Hierarchical Global–Local Planner for Zero-Shot Vision-and-Language Navigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current zero-shot Vision-and-Language Navigation (VLN) agents powered by MLLMs still tend to drift off course, halt prematurely, and achieve low overall success rates. We propose Three-Step Nav to counteract these failures with a three-view protocol: First, look forward to extract global landmarks and sketch a coarse plan. |
Wanrong Zheng; Yunhao Ge; Laurent Itti; |
| 130 | Active Measuring in Reinforcement Learning With Delayed Negative Effects Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formulate an AOMDP as a periodic partially observable MDP and propose an online RL algorithm based on belief states. |
Daiqi Gao; Ziping Xu; Aseel Rawashdeh; Predrag Klasnja; Susan Murphy; |
| 131 | Time-Aware Synthetic Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite its success across diverse applications, existing SC methods typically treat pre-intervention time indices as exchangeable, meaning they may fail to exploit temporal structure when strong trends are present. We propose Time-Aware Synthetic Control (TASC), a method that addresses this limitation by adopting a state-space model with a constant trend component while preserving the low-rank structure of the signal. |
Saeyoung Rho; Cyrus Illick; Samhitha Narasipura; Alberto Abadie; Daniel Hsu; Vishal Misra; |
| 132 | Sparse Offline Reinforcement Learning with Corruption Robustness Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In our setting, an adversary may arbitrarily perturb a fraction of the collected trajectories from a high-dimensional but sparse Markov decision process, and our goal is to estimate a near-optimal policy. |
Nam Phuong Tran; Andi Nika; Goran Radanovic; Long Tran-Thanh; Debmalya Mandal; |
| 133 | Structured Difference-of-Q Via Orthogonal Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a dynamic generalization of the R-learner (Nie et al 2021, Lewis and Syrgkanis 2021) for estimating and optimizing the difference of $Q^\pi$-functions, $Q^\pi(s,a)-Q^\pi(s,a_0)$, for potential discrete-valued actions $a,a_0$, which can be used to optimize multiple-valued actions without loss of generality. |
Defu Cao; Angela Zhou; |
| 134 | Train Less, Infer Faster: Efficient Model Finetuning and Compression Via Structured Sparsity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a scheme for effective finetuning via sparsification using training stochastic gates, which requires minimal trainable parameters, reduces inference time, and removes 20–40\% of model parameters without significant accuracy loss. |
Jonathan Svirsky; Yehonathan Refael; Ofir Lindenbaum; |
| 135 | High-Dimensional Analysis of Bootstrap Ensemble Classifiers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents a theoretical analysis of bootstrap techniques applied to the Least Square Support Vector Machine (LSSVM) ensemble in the context of large and growing sample sizes and feature dimensionalities. |
Malik Tiomoko; Hamza Cherkaoui; Mohamed El Amine Seddik; Cosme Louart; Ekkehard Schnoor; Balázs Kégl; |
| 136 | An Indicator of Membership Inference Security in Post-Training Quantized Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we investigate the impact of quantization procedures on privacy in data-driven models, focusing on their vulnerability to membership inference attacks. |
Eric AUBINAIS; Philippe Formont; Pablo Piantanida; Elisabeth Gassiat; |
| 137 | From Counts to Preferences: Preference-Driven Models for Spatio-Temporal Event Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a preference-driven framework that models event distributions through a two-stage “consider–then–choose” process: sparse gating captures limited attention, and utility functions guide selection within the consideration set. |
Chao Yang; Yiling Kuang; Shuang Li; |
| 138 | Provable Accelerated Bayesian Optimization with Knowledge Transfer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose the DeltaBO algorithm, which builds a novel uncertainty-quantification approach on the difference function $\delta$ between the source and target functions, which are allowed to belong to different Reproducing Kernel Hilbert Spaces (RKHSs). |
Haitao Lin; Boxin Zhao; Mladen Kolar; Chong Liu; |
| 139 | CAWI: Copula-Aligned Weight Initialization for Randomized Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To the best of our knowledge, this limitation remains unaddressed in the RdNN literature. To close this gap, we propose CAWI (Copula-Aligned Weight Initialization), a framework that draws input-to-hidden weights from a data-fitted copula that matches empirical dependence, ensuring the frozen projections respect inter-feature dependence without sacrificing closed-form solution. |
Mushir Akhtar; M. Tanveer; Mohd. Arshad; |
| 140 | When Can Federated Learning Match Centralized Learning? A PAC-Bayesian Generalization Gap Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The growing focus on distributed data and privacy has spurred the rise of Federated Learning (FL). Empirical studies show that, under equal resources, FL often underperforms centralized training, but the reasons behind this gap remain theoretically unclear. |
Xuanyu Chen; Shuai Wang; NAN YANG; Dong Yuan; |
| 141 | Provable Effects of Data Replay in Continual Learning: A Feature Learning Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present the first theoretical framework for analyzing full data-replay training in continual learning from a feature learning perspective. |
Meng Ding; Jinhui Xu; Kaiyi Ji; |
| 142 | Q-Learning with Shift-Aware Upper Confidence Bound in Non-Stationary Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While the Q-learning Upper Confidence Bound algorithm (QUCB) can discover a proper policy during learning, due to the distribution shifts, this policy can exploit sub-optimal rewards after the shift happens. To address this issue, we propose Density-QUCB (DQUCB), a shift-aware Q-learning UCB algorithm, which uses a transition density function to detect distribution shifts, then leverages its likelihood to enhance the uncertainty estimation quality of Q-learning UCB, resulting in a balance between exploration and exploitation. |
Ha Manh Bui; Felix Parker; Kimia Ghobadi; Anqi Liu; |
| 143 | A Recovery Theory for Diffusion Priors: Deterministic Analysis of The Implicit Prior Algorithm Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we develop a theoretical framework for analyzing deterministic diffusion-based algorithms for inverse problems, focusing on a deterministic version of the algorithm proposed by Kadkhodaie \& Simoncelli \cite{kadkhodaie2021stochastic}. |
Oscar Leong; Yann Traonmilin; |
| 144 | SQuaT: Self-Supervised Knowledge Distillation Via Student-Aware Quantized Teacher Features Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Knowledge Distillation (KD) is a common approach to address this challenge, but we observe that prior work combining QAT with KD suffers from a fundamental limitation: during distillation, the range mismatch between the teacher and the quantized student model induces an unattainable residual, resulting in an irreducible lower bound on the distillation loss. Motivated by this observation, we propose SQuaT (Student-Aware Quantized Teacher Features), a label-free QAT framework with KD that theoretically eliminates this lower bound by applying the student’s quantization parameters to quantize the teacher’s features during distillation. |
HyeonJun Lee; Hyeonsik Jo; Jinwoo Chung; Jangho Kim; |
| 145 | Dual Averaging Converges for Nonconvex Smooth Stochastic Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: On the other hand, in the non-convex setting, the situation is drastically different: While it is provable that SGD can minimize the gradient norm of non-convex smooth functions, no finite-time complexity guarantee for Stochastic Dual Averaging (SDA) was known in the same setting. In this paper, we close this gap by a reduction that views SDA as SGD applied to a sequence of implicitly regularized objectives. |
Tuo Liu; El Mehdi Saad; Wojciech Kotlowski; Francesco Orabona; |
| 146 | Beyond Johnson-Lindenstrauss: Uniform Bounds for Sketched Bilinear Forms Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, many modern analyses involve *sketched bilinear forms*, for which existing uniform bounds either do not apply or are not sharp on general sets. In this work, we develop a general framework to analyze such sketched bilinear forms, and derive uniform bounds in terms of geometric complexities of the associated sets. |
Rohan Deb; Qiaobo Li; Mayank Shrivastava; Arindam Banerjee; |
| 147 | Optimal Arm Elimination Algorithms for Combinatorial Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While extensions of upper confidence bound (UCB) algorithms arise naturally in this context, adapting arm elimination methods has proved more challenging. We introduce a novel elimination scheme that partitions arms into three categories (confirmed, active, and eliminated), and incorporates explicit exploration to update these sets. |
Yuxiao Wen; Yanjun Han; Zhengyuan Zhou; |
| 148 | The Good, The Bad, and The Sampled: A No-Regret Approach to Safe Online Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a method that jointly estimates the logistic parameter $\theta^{\star}$ and the feature distribution, using a conservative threshold on the logistic score to decide when to test. |
Tavor Baharav; Spyros Dragazis; Aldo Pacchiano; |
| 149 | Towards Motion-aware Referring Image Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For comprehensive evaluation, we introduce a new test split focusing on motion-centric queries, and introduce a new benchmark called M-Bench, where objects are distinguished primarily by actions. |
Chaeyun Kim; Seunghoon Yi; Yejin Kim; Yohan Jo; Joonseok Lee; |
| 150 | Balanced and Robust Multi-Treatment Experimental Designs Via Randomized Differencing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce GKK+, a new design for multi-arm randomized controlled trials. |
Qing Chen; Jing Jia; Peng Zhang; |
| 151 | Why Is Prompting Hard? Understanding Prompts on Binary Sequence Predictors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Taken together, this work takes an initial step towards understanding optimal prompts, from a statistical and empirical perspective that complements research on frontier models. |
Li Kevin Wenliang; Anian Ruoss; Jordi Grau-Moya; Marcus Hutter; Tim Genewein; |
| 152 | Non-Stationary Functional Bilevel Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **SmoothFBO**, the first algorithm for non-stationary FBO with both theoretical guarantees and practical scalability. |
Jason Bohne; Ieva Petrulionytė; Michael Arbel; Julien Mairal; Pawel Polak; |
| 153 | Regularized $f$-Divergence Kernel Tests Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a framework to construct practical kernel-based two-sample tests from the family of $f$-divergences. |
Mónica Ribero; Antonin Schrab; Arthur Gretton; |
| 154 | SPIRE: Conditional Personalization for Federated Diffusion Generative Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To enable personalized diffusion generative models, we propose Shared‑backbone Personal Identity Representation Embeddings (SPIRE), a framework that casts per‑client diffusion based generation as conditional generation in FL. |
Kaan Ozkara; Ruida Zhou; Suhas Diggavi; |
| 155 | Duality-based Residual Estimation for Fully Offline Value-based Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose the **Duality-based Residual Estimator (DRE)**, a simple offline validation metric for value-based offline RL. |
Kohei Miyaguchi; |
| 156 | Spectral Clustering for Directed Graphs Via Likelihood Estimation on Stochastic Block Models Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: In this paper, we leverage statistical inference on stochastic block models to guide the development of a spectral clustering algorithm for directed graphs. |
Ning Zhang; Xiaowen Dong; Mihai Cucuringu; |
| 157 | Linear Reasoning Vs. Proof By Cases: Obstacles for Large Language Models in FOL Problem Solving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our experimental results over leading LLMs demonstrate a substantial performance gap between linear reasoning and case-based reasoning problems. To further investigate this phenomenon, we provide a theoretical analysis grounded in graphical model, which provides an explanation for the observed disparity between the two types of reasoning problems. |
Yuliang Ji; Fuchen Shen; Jian Wu; Qiujie Xie; Yue Zhang; |
| 158 | In-memory Training on Analog Devices with Limited Conductance States Via Multi-tile Residual Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In practice, many promising memristive devices such as ReRAM offer only about 4-bit resolution due to fabrication constraints, and this limited update precision substantially degrades training accuracy. To enable on-chip training with these limited-state devices, this paper proposes a \emph{multi-tile residual learning} framework that sequentially learns on multiple crossbar tiles to compensate the residual errors from low-precision weight updates. |
Jindan Li; Zhaoxian Wu; Gaowen Liu; Tayfun Gokmen; Tianyi Chen; |
| 159 | Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper puts forth a minimal mathematical model that provably identifies when ICL can generalize out-of-distribution (OOD). |
Soo Min Kwon; Alec S. Xu; Can Yaras; Laura Balzano; Qing Qu; |
| 160 | Multiclass Local Calibration with The Jensen-Shannon Distance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches to multiclass calibration lack a notion of distance among inputs, which makes them vulnerable to proximity bias: predictions in sparse regions of the feature space are systematically miscalibrated. In this work, we address this main shortcoming by introducing a local perspective on multiclass calibration. |
Cesare Barbera; Lorenzo Perini; Giovanni De Toni; Andrea Passerini; Andrea Pugnana; |
| 161 | Learning Linear Regression with Low-Rank Tasks In-Context Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: It is particularly mysterious how ICL operates in real-world applications where tasks have a structure. In this work, we address this problem by analyzing a linear attention model trained on low-rank regression tasks. |
Kaito Takanami; Takashi Takahashi; Yoshiyuki Kabashima; |
| 162 | Learning with Incomplete Context: Linear Contextual Bandits with Pretrained Imputation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose PULSE-UCB, an algorithm that leverages pretrained models trained on the auxiliary data to impute missing features during online decision-making. |
Hao Yan; Heyan Zhang; Yongyi Guo; |
| 163 | Graph Learning Is Suboptimal in Causal Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We further analyze both the known and unknown parent set size regimes, establish novel regret lower bounds that capture the combinatorial structure of the action space. Building on these insights, we propose nearly optimal algorithms that bypass graph and parent recovery, demonstrating that parent identification is indeed unnecessary for regret minimization. |
Mohammad Shahverdikondori; Jalal Etesami; Negar Kiyavash; |
| 164 | Beyond Real Data: Synthetic Data Through The Lens of Regularization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present a learning-theoretic framework to quantify the trade-off between synthetic and real data. |
Amitis Shidani; Tyler Farghly; Yang SUN; Habib Ganjgahi; George Deligiannidis; |
| 165 | Counterfactually Fair Conformal Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through symmetrization of conformity scores across protected-attribute interventions, we prove that CF-CP results in counterfactually fair prediction sets while maintaining the marginal coverage property. |
Ozgur Guldogan; Neeraj Sarna; Yuanyuan Li; Michael Berger; |
| 166 | The Role of Causal Features in Strategic Classification for Robustness and Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: In strategic classification, an institution (e.g., a bank) anticipates adaptation from users who change their features to increase utility in a classification task (e.g., loan … |
António Góis; Sophia Günlük; Nir Rosenfeld; Nidhi Hegde; Simon Lacoste-Julien; Dhanya Sridhar; |
| 167 | Tight Regret Upper and Lower Bounds for Optimistic Hedge in Two-Player Zero-Sum Games Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study investigates the optimality of the dependence on $m$ and $n$ in the regret of optimistic Hedge. |
Taira Tsuchiya; |
| 168 | Evaluation of Large Language Models Via Coupled Token Generation Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: Consequently, a model may respond differently to the same prompt if asked multiple times. In this work, we argue that the evaluation and ranking of large language models should control for this randomization. |
Nina L. Corvelo Benz; Stratis Tsirtsis; Eleni Straitouri; Ivi Chatzi; Ander Artola Velasco; Suhas Thejaswi; Manuel Gomez Rodriguez; |
| 169 | On Kernel Based Variational Autoencoders Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we bridge Variational Autoencoders (VAEs) and kernel density estimations (KDEs) by approximating the posterior by the expectation of kernel density estimator and deriving a new lower bound of empirical log likelihood. |
Tian Qin; Wei-Min Huang; |
| 170 | Adaptive Replay Buffer for Offline-to-Online Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Standard methods, often relying on a fixed data-mixing ratio, struggle to manage the trade-off between early learning stability and asymptotic performance. To overcome this, we introduce the Adaptive Replay Buffer (ARB), a novel approach that dynamically prioritizes data sampling based on a lightweight metric we call ‘on-policyness’. |
Chihyeon Song; Jaewoo Lee; Jinkyoo Park; |
| 171 | ReTrack: Data Unlearning in Diffusion Models Through Redirecting The Denoising Trajectory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ReTrack, a fast and effective data unlearning method for diffusion models. |
Qitan Shi; Cheng Jin; Jiawei Zhang; Yuantao Gu; |
| 172 | Differentially Private Algorithms for The Stochastic Compositional Optimization Problem Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study the stochastic compositional optimization problem under the constraint of differential privacy. |
Zhuanghua Liu; Weida Li; Xiaokui Xiao; Bryan Kian Hsiang Low; |
| 173 | Sequential 1-bit Mean Estimation with Near-Optimal Sample Complexity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study the problem of distributed mean estimation with 1-bit communication constraints. |
Ivan Lau; Jonathan Scarlett; |
| 174 | Neuron Block Dynamics for XOR Classification with Zero-Margin Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast, we study zero-margin nonlinear classification by analyzing the Gaussian XOR problem, where inputs are Gaussian and the XOR decision boundary determines labels. |
Guillaume Braun; Masaaki Imaizumi; |
| 175 | Simplex-to-Euclidean Bijections for Categorical Flow Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a method for learning and sampling from probability distributions supported on the simplex. |
Bernardo Williams; Victor M. Yeom-Song; Marcelo Hartmann; Arto Klami; |
| 176 | Personalized Incentive Alignment: Correcting Utility-Driven Selection Bias in A/B Tests Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we posit that participants act to maximize individual incentives. |
Jiachun Li; Yang Meng; David Simchi-Levi; Chonghuan Wang; |
| 177 | ConDiSim: Conditional Diffusion Models for Simulation-Based Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present ConDiSim, a conditional diffusion model for simulation-based inference in complex systems with intractable likelihoods. |
Mayank Nautiyal; Andreas Hellander; Prashant Singh; |
| 178 | Complexity-Aware Deep Symbolic Regression with Robust Risk-Seeking Policy Gradients Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel deep symbolic regression (DSR) approach to enhance the robustness and interpretability of data-driven mathematical expression discovery. |
Zachary Bastiani; Mike Kirby; Jacob Hochhalter; Shandian Zhe; |
| 179 | Batch-Adaptive Causal Annotations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While ground-truth outcomes can sometimes be obtained through costly data annotation or follow-up, budget constraints typically allow only a fraction of the dataset to be labeled. We address this challenge by optimizing which data points should be sampled for outcome information in order to improve efficiency in average treatment effect estimation with missing outcomes. |
Ezinne Nwankwo; Lauri Goldkind; Angela Zhou; |
| 180 | Efficient Flow Matching Using Latent Variables Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we present $\texttt{Latent-CFM}$, which provides efficient training strategies by conditioning on the features extracted from data using pretrained deep latent variable models. |
Anirban Samaddar; Yixuan Sun; Viktor Nilsson; Sandeep Madireddy; |
| 181 | Formally Exploring Time-Series Anomaly Detection Evaluation Metrics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we formalize the problem by introducing verifiable properties of evaluation metrics that individually reflect important aspects of anomaly detection in time series. |
Dennis Wagner; Arjun Nair; Billy Joe Franks; Justus Arweiler; Aparna Muraleedharan; Indra Jungjohann; Fabian Hartung; Andriy Balinskyy; Saurabh Varshneya; Mayank Chetan Ahuja; Nabeel Hussain Syed; Mayank Nagda; Philipp Liznerski; Steffen Reithermann; Maja Rudolph; Sebastian Josef Vollmer; Ralf Schulz; Torsten Katz; Stephan Mandt; Michael Bortz; Heike Leitte; Daniel Neider; Jakob Burger; Fabian Jirasek; Hans Hasse; Sophie Fellenz; Marius Kloft; |
| 182 | Learning How Deep to Go: Self-Scaling Deep Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, these architectures incur substantial computational and energy costs, and selecting the optimal network depth in advance remains an open challenge. In this paper, we introduce SCALE-RL, a self-scaling DRL framework that dynamically adjusts its architectural depth during training, allowing the network to automatically adapt its depth to the task. |
Michelangelo Vegliò; Marco Fantozzi; Antonio Di Cecco; Carlo Metta; Flora Angileri; Simone Treccani; Adrienne Chloe Rayos Macazar; Silvia Giulia Galfre’; Maurizio Parton; Francesco Morandin; |
| 183 | Local Regression on Path Spaces with Signature Metrics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a functional Nadaraya-Watson estimator that combines the signature transform from rough path theory with local kernel regression. |
Christian Bayer; Davit Gogolashvili; Luca Pelizzari; |
| 184 | Efficient Logistic Regression with Mixture of Sigmoids Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper studies the Exponential Weights (EW) algorithm with an isotropic Gaussian prior for online logistic regression. |
Federico Di Gennaro; Saptarshi Chakraborty; Nikita Zhivotovskiy; |
| 185 | Busemann Functions in The Wasserstein Space: Existence, Closed-Forms, and Applications to Slicing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we investigate the existence and computation of Busemann functions in Wasserstein space, which admits geodesic rays. |
Clément Bonet; Elsa Cazelles; Lucas Drumetz; Nicolas Courty; |
| 186 | NeST-BO: Fast Local Bayesian Optimization Via Newton-Step Targeting of Gradient and Hessian Information Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose NeST-BO, a curvature-aware local BO method that targets a (modified) Newton step by jointly learning gradient and Hessian information with Gaussian process (GP) surrogates, and selecting evaluations via a one-step lookahead bound on the Newton-step error. |
Wei-Ting Tang; Akshay Kudva; Joel Paulson; |
| 187 | Split-Flows: Measure Transport and Information Loss Across Molecular Resolutions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce split-flows, a novel flow-based approach that reinterprets backmapping as a continuous-time measure transport across resolutions. |
Sander Hummerich; Ullrich Koethe; Tristan Bereau; |
| 188 | Beyond Black-Box Predictions: Identifying Marginal Feature Effects in Tabular Transformer Networks Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: To bridge the gap between intelligibility and performance, we propose an adaptation of tabular transformer networks designed to identify marginal feature effects. |
Anton Frederik Thielmann; Arik Reuter; Benjamin Säfken; |
| 189 | Explanation Design in Strategic Learning: Sufficient Explanations That Induce Non-harmful Responses Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we analyse widely used explanation methods and establish a necessary condition to prevent explanations from inducing self-harming responses. |
Kiet Q. H. Vo; Siu Lun Chau; Masahiro Kato; Yixin Wang; Krikamol Muandet; |
| 190 | Amortized Safe Active Learning for Real-Time Data Acquisition: Pretrained Neural Policies from Simulated Nonparametric Functions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose amortized AL for regression and amortized safe AL, replacing expensive online computations with a pretrained neural policy. |
Cen-You Li; Marc Toussaint; Barbara Rakitsch; Christoph Zimmer; |
| 191 | Scalable Spatiotemporal Inference with Biased Scan Attention Transformer Neural Processes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These applications have placed increasing pressure on the scalability of these models, with many architectures compromising accuracy for scalability. In this paper, we demonstrate that this trade-off is often unnecessary, particularly when modeling fully or partially translation-invariant processes. |
Daniel Jenson; Jhonathan Navott; Piotr Grynfelder; Mengyan Zhang; Makkunda Sharma; Elizaveta Semenova; Seth Flaxman; |
| 192 | Conformal Margin Risk Minimization: An Envelope Framework for Robust Learning Under Label Noise Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ***C**onformal **M**argin **R**isk **M**inimization (CMRM)*, a plug-and-play envelope framework that improves \emph{any} classification loss under label noise by adding a single quantile-calibrated regularization term, with no privileged knowledge or training pipeline modification. |
Yuanjie Shi; Peihong Li; Zijian Zhang; Jana Doppa; Yan Yan; |
| 193 | Improving Adaptive Moment Optimization Via Preconditioner Diagonalization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the gradient statistics employed by these methods often do not leverage sufficient gradient covariance information, leading to suboptimal updates in certain directions of the parameter space and potentially slower convergence. In this work, we keep track of such covariance statistics in the form of a structured preconditioner matrix. |
Son Nguyen; Bo Liu; Lizhang Chen; qiang liu; |
| 194 | Slithering Through Gaps: Capturing Discrete Isolated Modes Via Logistic Bridging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This limitation makes it difficult for sampling methods to achieve adequate mixing and convergence in high-dimensional multimodal discrete spaces. To address these challenges, we propose Hyperbolic Secant-squared Gibbs-Sampling (HiSS), a novel family of sampling algorithms that integrates a Metropolis-within-Gibbs framework to enhance mixing efficiency. |
Pinaki Mohanty; Ruqi Zhang; |
| 195 | One-Step Diffusion Samplers Via Self-Distillation and Deterministic Flow Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce one-step diffusion samplers which learn a step-conditioned ODE so that one large step reproduces the trajectory of many small ones via a state-space consistency loss. |
Pascal Jutras Dube; Jiaru Zhang; Ziran Wang; Ruqi Zhang; |
| 196 | Integrating Feature Correlation in Differential Privacy with Applications in DP-ERM Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Standard differential privacy imposes uniform privacy constraints across all features, overlooking the inherent distinction between sensitive and insensitive features in practice. In this paper, we introduce a relaxed definition of differential privacy that accounts for such heterogeneity, allowing certain features to be treated as insensitive even when correlated with sensitive ones. |
Tianyu Wang; Luhao Zhang; Rachel Cummings; |
| 197 | FocusViT: Faithful Explanations for Vision Transformers Via Gradient-Guided Layer-Skipping Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose FocusViT, a novel explainability framework that integrates gradient-weighted attention attribution with validation-based, faithfulness-driven layer aggregation. |
Mohsin Ali; Haider Raza; John Q Gan; Muhammad Haris Khan; |
| 198 | Corruption Robust Thompson Sampling for Gaussian Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The main challenge is that one can no longer compute the actual posteriors for the true reward, as the agent can only observe the rewards after corruption. In this work, we solve this problem by computing pseudo-posteriors that are less likely to be manipulated by the attack. |
Yinglun Xu; Zhiwei Wang; Gagandeep Singh; |
| 199 | SetPINNs: Set-based Physics-informed Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SetPINNs, a framework that effectively captures local dependencies. |
Mayank Nagda; Phil Ostheimer; Thomas Specht; Frank Rhein; Fabian Jirasek; Stephan Mandt; Marius Kloft; Sophie Fellenz; |
| 200 | An Illusion of Unlearning? Assessing Machine Unlearning Through Internal Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Assuming neural collapse in the original model, we further demonstrate that adjusting only the classifier can achieve negligible forget accuracy while preserving retain accuracy, and we corroborate this with experiments using classifier-only fine-tuning. Motivated by these findings, we propose MU methods based on a class-mean features (CMF) classifier, which explicitly enforces alignment between features and classifiers. |
Yichen Gao; Altay Unal; Akshay Rangamani; Zhihui Zhu; |
| 201 | Precise Dynamics of Diagonal Linear Networks: A Unifying Analysis By Dynamical Mean-Field Theory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present a unified analysis of various phenomena in the gradient flow dynamics of DLNs. |
Sota Nishiyama; Masaaki Imaizumi; |
| 202 | HGT-FD: Hypergraph Transformer for Fraud Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose HGT-FD for hypergraph-based fraud detection. |
Yintao Cai; Yunjiong Liu; Zelong Yang; Shuyang Fang; Xiaoping Min; |
| 203 | Structural Alignment Improves Graph Test-Time Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Test-Time Structural Alignment (TSA), a novel algorithm for Graph Test-Time Adaptation (GTTA) that adapts a pretrained model to align graph structures during inference without the cost of retraining. |
Hans Hao-Hsun Hsu; Shikun Liu; Han Zhao; Pan Li; |
| 204 | Learning Geometry and Topology Via Multi-Chart Flows Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we first propose the general training scheme for learning such a collection of flows, and secondly we develop the first numerical algorithms for computing geodesics on such manifolds. Empirically, we demonstrate that this leads to highly significant improvements in topology estimation. |
Hanlin Yu; Søren Hauberg; Marcelo Hartmann; Arto Klami; Georgios Arvanitidis; |
| 205 | Orthogonal Representation Learning for Estimating Causal Quantities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: and (2) Can a balancing constraint — commonly proposed technique in the representation learning literature — provide improvements to Neyman-orthogonality? We address these two questions through our theoretical and empirical analysis, where we introduce a unifying framework that connects representation learning with Neyman-orthogonal learners (namely, OR-learners). |
Valentyn Melnychuk; Dennis Frauen; Jonas Schweisthal; Stefan Feuerriegel; |
| 206 | Gradient-Flow SDEs Have Unique Transient Population Dynamics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We dispel the need for this assumption by providing a complete characterization of identifiability: the gradient-flow drift and Brownian diffusivity are jointly identifiable from temporal marginals if and only if the process is observed outside of equilibrium. Given this fundamental result, we propose nn-APPEX, the first Schrödinger Bridge–based inference method that can simultaneously learn the drift and diffusion of a gradient-flow SDE solely from observed marginals. |
Vincent Guan; Joseph Janssen; Nicolas Lanzetti; Antonio Terpin; Geoffrey Schiebinger; Elina Robeva; |
| 207 | Unified Causal Discovery and Missing Data Imputation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose LOGIC, a framework that performs causal discovery and causally consistent imputation jointly. |
Osman Mian; Jens Kleesiek; Michael Kamp; |
| 208 | On Barycenter Computation: Analyzing Semi-Unbalanced Optimal Transport-based Method on Bures-Wasserstein Manifold Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop optimization algorithms on Bures-Wasserstein manifold, named the Exact Geodesic Gradient Descent and Hybrid Gradient Descent algorithms. |
Ngoc-Hai Nguyen; Le Quang Dung; Hoang-Phi Nguyen; Tung Pham; Nhat Ho; |
| 209 | The Minimax Lower Bound of Kernel Stein Discrepancy Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To the best of our knowledge, all existing KSD estimators with known rate achieve $\sqrt n$-convergence. In this work, we present two complementary results (with different proof strategies), establishing that the minimax lower bound of KSD estimation is $n^{-1/2}$ and settling the optimality of these estimators. |
Jose Cribeiro-Ramallo; Agnideep Aich; Florian Kalinke; ASHIT BARAN AICH; Zoltán Szabó; |
| 210 | Stationarity-Aware Causal Discovery in Time Series Via Minimal Separating Sets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We observe that the stationary graph structure and autoregressive edges impose many meaningful constraints on the separating sets between variables at different time lags. After characterizing the behavior of such separating sets, we propose a novel causal discovery algorithm that exploits this structure of minimal separating sets. |
Shanyun Gao; Raghavendra Addanki; Tong Yu; Ryan A. Rossi; Qifan Song; Murat Kocaoglu; |
| 211 | On Relation-Aware Slicing in Cross-Domain Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a remedy, we propose an optimization-free slicing distribution that provides fast sampling for the Monte Carlo approximation. |
Dhruv Sarkar; Aprameyo Chakrabartty; Anish Chakrabarty; Swagatam Das; |
| 212 | On The Finite-Sample Bias of Minimizing Expected Wasserstein Loss Between Empirical Distributions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We derive closed-form expressions for the expected Wasserstein loss in one dimension and, focusing on location–scale models, provide an analytic characterization of the bias. This analysis reveals that finite-sample bias occurs whenever the expected loss varies along the diagonal subspace where parameter values coincide, and we propose a simple correction scheme that removes this effect. |
Cheongjae Jang; Yung-Kyun Noh; |
| 213 | Near-Optimal Dropout-Robust Sortiton Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our main contribution is an efficient loss-minimizing algorithm, which remains optimal as we vary the maximizer’s power from worst-case to average case. |
Maya Pal Gambhir; Bailey Flanigan; Aaron Roth; |
| 214 | Improving Semantic Uncertainty Quantification in Language Model Question-Answering Via Token-Level Temperature Scaling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We demonstrate that optimising a single scalar temperature, which, we argue, provides a suitable inductive bias, is a surprisingly simple yet effective solution. |
Tom A. Lamb; Desi R. Ivanova; Philip Torr; Tim G. J. Rudner; |
| 215 | Multi-Metric Adaptive Experimental Design Under A Fixed Budget with Validation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes a fixed-budget multi-metric AED framework with a two-phase structure: an adaptive exploration phase to identify the best treatment, and a validation phase with an A/B test to verify the treatment’s quality and infer statistics. |
Qining Zhang; Tanner Fiez; Yi Liu; Wenyang Liu; |
| 216 | TENDE: Transfer Entropy Neural Diffusion Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing estimation methods suffer from the curse of dimensionality, require restrictive distributional assumptions, or need exponentially large datasets for reliable convergence. We address these limitations in the literature by proposing TENDE (Transfer Entropy Neural Diffusion Estimation), a novel approach that leverages score-based diffusion models to estimate transfer entropy through conditional mutual information. |
Simon Pedro Galeano Munoz; Maurizio Filippone; Giulio Franzese; Mustapha Bounoua; Pietro Michiardi; |
| 217 | TexTSC: Class-Texture Preserving Data Condensation for Time Series Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present TexTSC, a condensation framework that preserves class structure using spectro-temporal second-order statistics instead of trajectory replay. |
Pouya Hosseinzadeh; Peiyu Li; Omar Bahri; Soukaina Filali Boubrahimi; Shah Muhammad Hamdi; |
| 218 | Dyno-Net: A Dynamic Feature Extraction Model for Gastrointestinal Polyp Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Dyno-Net, a dynamic feature extraction framework integrating multi-scale fusion (DynoFPN), adaptive convolution (DynoConv), and boundary refinement (RefineDet\_LSCSBD), achieving 23.5\% higher fusion efficiency, 17.8\% better detection of small/atypical polyps, and mean IoU improvement from 0.68 to 0.81. |
Zijie Song; Jingjing Wan; Xianchun Meng; Qingye Hua; Wenjie Zhu; Bolun Chen; WEI SHAO; |
| 219 | Efficient Bilevel Optimization with KFAC-Based Hypergradients Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We build on implicit function theorem-based algorithms and propose to incorporate Kronecker-factored approximate curvature (KFAC), yielding curvature-aware hypergradients with a better performance efficiency trade-off than Conjugate Gradient (CG) or Neumann methods and consistently outperforming unrolling. |
Disen Liao; Felix Dangel; Yaoliang Yu; |
| 220 | Convergence of Projected Stochastic Natural Gradient Variational Inference for Various Step Size and Sample or Batch Size Schedules Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we define and study a projected stochastic NGVI when variational distributions form an exponential family. |
Thomas Guilmeau; Hadrien Hendrikx; Florence Forbes; |
| 221 | Nearly Optimal Best Arm Identification for Semiparametric Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study fixed-confidence Best Arm Identification (BAI) in semiparametric bandits, where rewards are linear in arm features plus an unknown additive baseline shift. |
Seok-Jin Kim; |
| 222 | Causal Additive Models with Unobserved Causal Paths and Backdoor Paths Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These conditions rely on new characterizations of regression sets to determine independence among regression residuals and conditional independencies among observed variables. Building on these results, we introduce a search algorithm that incorporates these innovations and prove its soundness and completeness. |
Thong Pham; Takashi Nicholas Maeda; Shohei Shimizu; |
| 223 | On The Neural Feature Ansatz for Deep Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Assuming gradient flow dynamics with balanced weight initialization, the NFA was proven to hold throughout training for two-layer linear networks with exponent $\alpha = 1/2$ (Radhakrishnan et al., 2024). We extend this result to networks with $L \geq 2$ layers, showing that the NFA holds with exponent $\alpha = 1/L$, thus demonstrating a depth dependency of the NFA. |
Edward Tansley; Estelle Massart; Coralia Cartis; |
| 224 | Kernel Treatment Effects with Adaptively Collected Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present the first kernel-based framework for distributional inference under adaptive data collection. |
Houssam Zenati; Bariscan Bozkurt; Arthur Gretton; |
| 225 | Statistical-computational Gap in Multiple Gaussian Graph Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate the existence of a statistical-computational gap in multiple Gaussian graph alignment. |
Bertrand Even; Luca Ganassali; |
| 226 | Hyperbolic Part-Whole Image Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce a hyperbolic prototypical segmentation framework capable of simultaneously representing multiple granularity levels within a unified embedding space. |
Mikhail Vlasenko; Mina Ghadimi Atigh; Pascal Mettes; |
| 227 | Statistical Inference for Explainable Boosting Machines Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We provide an alternative using recent advances in statistical inference for gradient boosting, deriving methods for statistical inference as well as end-to-end theoretical guarantees. |
Haimo Fang; Kevin Tan; Jonathan Pipping; Giles Hooker; |
| 228 | Incoherence in Goal-Conditioned Autoregressive Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We prove that it decreases incoherence and leads to an improvement in return, and we aim to characterise the resulting trajectory of policies. |
Jacek Karwowski; Raymond Douglas; |
| 229 | MDPs with A State Sensing Cost Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formulate this as an expected discounted cost Markov Decision Process (MDP), wherein the agent incurs an additional cost for sensing its next state, but has the option to take actions while remaining `blind’ to the system state. We pose this problem as a classical discounted cost MDP with an expanded (countably infinite) state space. |
Vansh Kapoor; Jayakrishnan Nair; |
| 230 | Sparse Linear Bandits with Blocking Constraints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate the high-dimensional sparse linear bandits problem in a data-poor regime where the time horizon is much smaller than the ambient dimension and number of arms. |
Adit Jain; Soumyabrata Pal; Sunav Choudhary; Ramasuri Narayanam; Harshita Chopra; Vikram Krishnamurthy; |
| 231 | Stochastic Bandits on Mixture Distributions: Metrics & Regret Bounds Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We consider ranking arms using functions of the mixture parameters and propose methods to minimize the cumulative regret with respect to the induced ranking. |
Adit Jain; Sujay Bhatt; Alec Koppel; |
| 232 | Regret Guarantees for Linear Contextual Stochastic Shortest Path Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose LR-CSSP, an algorithm that achieves a regret bound of $\widetilde{O}(K^{2/3} d^{2/3} |S| |A|^{1/3} B_\star^2 T_\star \log (1/ \delta))$, where $K$ is the number of episodes, $d$ is the context dimension, $S$ and $A$ are the sets of states and actions respectively, $B_\star$ bounds the optimal cumulative loss and $T_\star$, unknown to the learner, bounds the expected time for the optimal policy to reach the goal. |
Dor Polikar; Alon Cohen; |
| 233 | Nonparametric Multi Change Point Detection for Markov Chains Via Adaptive Clustering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite this broad applicability across economics, computer science, and planetary sciences, rigorous, nonparametric techniques for change point detection with non-independent and identically distributed (i.i.d.) datasets has remained elusive. This paper establishes such guarantees by proposing a non-parametric clustering algorithm which can accurately obtain the change points from a given Markovian dataset of length $n$. |
Imon Banerjee; Jiaqi Lei; Sanjay Mehrotra; |
| 234 | Randomized HyperSteiner: A Stochastic Delaunay Triangulation Heuristic for The Hyperbolic Steiner Minimal Tree Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Randomized HyperSteiner (RHS), a stochastic Delaunay triangulation heuristic that incorporates randomness into the expansion process and refines candidate trees via Riemannian gradient descent optimization. |
Aniss Aiman Medbouhi; Alejandro García-Castellanos; Giovanni Luca Marchetti; Daniel Pelt; Erik J Bekkers; Danica Kragic; |
| 235 | Robust Generalization with Adaptive Optimal Transport Priors for Decision-Focused Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing Sinkhorn Distributionally Robust Optimization (DRO) methods provide theoretical guarantees but rely on a fixed reference distribution, which limits their adaptability. We propose a Prototype-Guided Distributionally Robust Optimization (PG-DRO) framework that learns class-adaptive priors from abundant base data via hierarchical optimal transport and embeds them into the Sinkhorn DRO formulation. |
Haixiang Sun; Andrew Liu; |
| 236 | Entropic Projection Alignment: Estimating, Explaining, and Improving Model Performance Under Distribution Shift Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a unified framework for addressing three key challenges of distribution shift: (1) estimating a model’s performance on an unlabeled target domain, (2) explaining the shift by identifying the features responsible, and (3) improving the target domain performance. |
Salim I. Amoukou; Emanuele Albini; Tom Bewley; Saumitra Mishra; Manuela Veloso; |
| 237 | An Evaluation of Cost Functions for Algorithmic Recourse Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose four metrics to evaluate whether currently popular cost functions in recourse satisfy the minimal requirements for meaningful distance calculations. |
Eoin M. Kenny; Allan Anzagira; Tom Bewley; Freddy Lecue; Manuela Veloso; |
| 238 | Archetypal Graph Generative Models: Explainable and Identifiable Communities Via Anchor-Dominant Convex Hulls Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present GraphHull, an explainable generative model that represents networks using two levels of convex hulls. |
Nikolaos Nakis; Chrysoula Kosma; Panagiotis Promponas; Michail Chatzianastasis; Giannis Nikolentzos; |
| 239 | Process-Tensor Tomography of SGD: Measuring Non-Markovian Memory Via Back-Flow of Distinguishability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We model neural training as a classical multi-time map from controllable interventions—batch choices, augmentations, and optimizer micro-steps—to model predictions on a fixed probe set. On this basis, we introduce a simple, model-agnostic witness of training memory based on back-flow of distinguishability. |
Vasileios Sevetlidis; George Pavlidis; |
| 240 | Projection-free Algorithms for Online Convex Optimization with Adversarial Constraints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our algorithm is conceptually simple. It constructs a surrogate cost function as a nonnegative linear combination of the cost and constraint functions, and feeds these surrogate costs into a novel adaptive online conditional gradient subroutine introduced in this paper. |
Dhruv Sarkar; Aprameyo Chakrabartty; Subhamon Supantha; Palash Dey; Abhishek Sinha; |
| 241 | On The Number of Conditional Independence Tests in Constraint-based Causal Discovery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we establish an algorithm that achieves a better complexity of $p^{\mathcal{O}(s)}$ tests, where $p$ is the number of nodes in the graph and $s$ denotes the maximum undirected clique size of the underlying essential graph. |
Marc Franquesa Monés; Jiaqi Zhang; Caroline Uhler; |
| 242 | Learning When Not to Learn: Risk-Sensitive Abstention in Bandits with Unbounded Rewards Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we formalize a model of learning with unbounded rewards without a mentor as a two-action contextual bandit with an abstain option: at each round the agent observes an input and chooses either to abstain (always 0 reward) or to commit (execute a preexisting task policy). |
Sarah Liaw; Benjamin Plaut; |
| 243 | Accelerated Distributed Optimization with Compression and Error Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel algorithm, ADEF (**A**ccelerated **D**istributed **E**rror **F**eedback), which integrates Nesterov acceleration, contractive compression, error feedback, and gradient difference compression. |
Yuan Gao; Anton Rodomanov; Jeremy Rack; Sebastian U Stich; |
| 244 | Explore-then-Commit for Nonstationary Linear Bandits with Latent Dynamics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an explore-then-commit algorithm for a finite horizon $T$. |
Sunmook Choi; Yahya Sattar; Yassir Jedra; Maryam Fazel; Sarah Dean; |
| 245 | Tyler’s M-estimator Through The Lens of Convex-Concave Programming Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In particular, no deterministic global convergence rate has been established, despite decades of research. In this paper, we resolve this longstanding question by interpreting the fixed-point iteration as an instance of the convex-concave procedure, and analyzing it through a novel combination of relative smoothness and Riemannian optimization. |
Daniel Cederberg; |
| 246 | Practical and Efficient Rashomon Set Sampling for Model Interpretability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These two axioms are not satisfied by most known attribution methods, which we consider to be a fundamental weakness. Building on these axioms, we propose an $\epsilon$-subgradient-based sampling framework and quantify effectiveness with *Search Efficiency Ratio* (SER) and *Functional Explanation Range* (FER). |
Sichao Li; Amanda S Barnard; Quanling Deng; |
| 247 | The Rashomon Effect for Visualizing High-Dimensional Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Dimension reduction (DR) is inherently non-unique: multiple embeddings can preserve the structure of high-dimensional data equally well while differing in layout or geometry. In this paper, we formally define the Rashomon set for DR—the collection of `good’ embeddings—and show how embracing this multiplicity leads to more powerful and trustworthy representations. |
Yiyang Sun; Haiyang Huang; Gaurav Rajesh Parikh; Cynthia Rudin; |
| 248 | Canopy Tree Height Estimation Using Quantile Regression: Modeling and Evaluating Uncertainty in Remote Sensing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we show that, with minor modifications of a given prediction head, existing models can be adapted to provide statistically calibrated uncertainty estimates via quantile regression. |
Karsten Schrödter; Jan Pauls; Fabian Gieseke; |
| 249 | Tight Lower Bounds and Optimal Algorithms for Stochastic Nonconvex Optimization with Heavy-Tailed Noise Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Beyond expectation guarantees, we introduce a new algorithm, Double-Clipped NSGD-MVR, which allows the derivation of high-probability convergence rates under weaker assumptions than in previous works. |
Adrien Fradin; Abdurakhmon Sadiev; Laurent Condat; Peter Richtárik; |
| 250 | $\epsilon$-Identifiability of Causal Quantities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper shows how approximate identifiability is still possible for several probabilities of causation. |
Ang Li; Scott Mueller; Xin Shu; Judea Pearl; |
| 251 | Differential Privacy in Kernelized Contextual Bandits Via Random Projections Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel algorithm that achieves the state-of-the-art cumulative regret of $\widetilde{\mathcal{O}}(\sqrt{\gamma_TT}+\frac{\gamma_T}{\varepsilon_{\text{DP}}})$ and $\widetilde{\mathcal{O}}(\sqrt{\gamma_TT}+\frac{\gamma_T\sqrt{T}}{\varepsilon_{\text{DP}}})$ over a time horizon of $T$ in the joint and local models of differential privacy, respectively, where $\gamma_T$ is the effective dimension of the kernel and $\varepsilon_{\text{DP}} > 0$ is the privacy parameter. |
Nikola Pavlovic; Sudeep Salgia; Qing Zhao; |
| 252 | RamPINN: Recovering Raman Spectra From Coherent Anti-Stokes Spectra Using Embedded Physics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose RamPINN, a model that learns to recover Raman spectra from given CARS spectra. |
Sai Karthikeya Vemuri; Adithya Ashok Chalain Valapil; Tim Büchner; Joachim Denzler; |
| 253 | GRANITE: A Generalized Regional Framework for Identifying Agreement in Feature-Based Explanations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose GRANITE, a generalized regional explanation framework that partitions the feature space into regions where interaction and distribution influences are minimized. |
Julia Herbinger; Gabriel Laberge; Maximilian Muschalik; Yann Pequignot; Marvin N. Wright; Fabian Fumagalli; |
| 254 | Policy-Oriented Binary Classification: Improving (KD-)CART Final Splits for Subpopulation Targeting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Maximizing Distance Final Split (MDFS), which generates split rules that strictly dominate CART/KD-CART under the unique intersect assumption. |
Bill Wang; Zhenbang Jiao; Fangyi Wang; |
| 255 | SiGHT: A Self-Supervised Graph-based Hallucination DeTection Framework for Domain-Specific LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge the gap, we propose SiGHT, a self-supervised graph framework designed for efficient hallucination detection in specialized contexts. |
Zi-Ying Chen; Meng-Fen Chiang; Wen-Chih Peng; |
| 256 | Active Measurement of Two-Point Correlations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study the problem of measuring the 2PCF over a large set of points, restricted to a subset satisfying a property of interest. |
Max Hamilton; Daniel Sheldon; Subhransu Maji; |
| 257 | Variational Grey-Box Dynamics Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast, physics-based simulation models described by ODEs/PDEs remain interpretable, but may have missing or unknown terms, unable to fully describe real-world observations. We bridge this gap with a novel grey-box method that integrates incomplete physics models directly into generative models. |
Gurjeet Sangra Singh; Frantzeska Lavda; Giangiacomo Mercatali; Alexandros Kalousis; |
| 258 | Local Causal Discovery for Statistically Efficient Causal Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Local Optimal Adjustments Discovery (LOAD), a sound and complete causal discovery approach that combines the computational efficiency of local methods with the statistical optimality of global methods. |
Mátyás Schubert; Tom Claassen; Sara Magliacane; |
| 259 | Identifiability of Potentially Degenerate Gaussian Mixture Models With Piecewise Affine Mixing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on our theoretical results, we propose a two-stage method to estimate the latent variables by enforcing sparsity and Gaussianity in the learned representations. |
Danru Xu; Sebastien Lachapelle; Sara Magliacane; |
| 260 | Fundamental Limits for Weighted Empirical Approximations of Exponentially Tilted Distributions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this article, we discuss the asymptotic efficiency of an estimator obtained by exponentially tilting the empirical distribution. |
Sarvesh Ravichandran Iyer; Himadri Mandal; Dhruman Gupta; Rushil Gupta; Agniv Bandyopadhyay; Achal Bassamboo; Sandeep Kumar Juneja; Varun Gupta; |
| 261 | On The Convergence and Stability of Distributed Sub-model Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, those random sampling of sub-models may not give satisfying convergence performance. In this paper, observing the success of SGD with shuffling, we propose a distributed shuffled sub-model training, where the full model is partitioned into several sub-models in advance, and the server shuffles those sub-models, sends each of them to clients at each round, and by the end of local updating period, clients send back the updated sub-models, and server averages them. |
Yuyang Deng; Fuli Qiao; Mehrdad Mahdavi; |
| 262 | Breaking Data Symmetry Is Needed For Generalization in Feature Learning Kernels Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Grokking occurs when a model achieves high training accuracy but generalization to unseen test points happens long after that. |
Marcel Tomàs Bernal; Neil Rohit Mallinar; Mikhail Belkin; |
| 263 | Variance Constrained Distribution Alignment in Few-shot Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such instability often leads to intra-class distribution drift and degraded generalization under few sample regimes. To address these challenges, we propose a method that can model class level latent distributions for flexible and efficient few shot synthesis. |
Xiaohong Cai; Yi SUN; Zhaowen Lin; Tianwei Cai; |
| 264 | Disentangling Federated Learning Heterogeneity: A Dual-Perspective Analysis of Quantifying Skew Versus Scarcity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Skew-Scarcity Disentanglement Indicator (SSDI), a novel metric that decomposes heterogeneity into two disentangled components: Label Distribution Skew (LDS) (quantity skew of present labels) and Label Coverage Deficiency (LCD) (deviation due to missing labels). |
Wenkai Zeng; NAN YANG; Zhiyu Zhu; Zhibo Jin; Dong Yuan; |
| 265 | ECAI: Efficient Convolution Activation Inversion for Constant-Memory Convolutional Neural Networks Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel approach that achieves constant activation memory usage during the training of convolutional neural networks (CNNs), addressing a key memory bottleneck in the backward pass. |
Changhyeon Lee; Seulki Lee; |
| 266 | EventFlow: Forecasting Temporal Point Processes with Flow Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose EventFlow, a non-autoregressive generative model for temporal point processes. |
Gavin Kerrigan; Kai Nelson; Padhraic Smyth; |
| 267 | Beyond Binning: Soft Task Reformulation for Deep Regression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a novel method aimed at improving test-time performance of neural networks on regression tasks. |
Lawrence Stewart; Francis Bach; Quentin Berthet; |
| 268 | Bayesian Fourier Features for Reduced Rank Gaussian Processes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel Fourier feature approach leveraging Bayesian quadrature methods to construct reduced-rank approximations of the Gaussian process kernel. |
Cristian A. Galvis-Florez; George Whittle; Michael A Osborne; Simo Särkkä; |
| 269 | DRAUN: An Optimization-Agnostic Data Reconstruction Attack on Federated Unlearning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work presents DRAUN, the first attack framework to reconstruct unlearned data in FU systems. |
Hithem Lamri; Manaar Alam; Haiyan Jiang; Michail Maniatakos; |
| 270 | Robust Learning of A Group DRO Neuron Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our goal is to identify a best-fit neuron parameterized by ${\boldsymbol w}\_{\*}$ that performs well under the most challenging reweighting of the groups. |
Guyang Cao; Shuyao Li; Sushrut Karmalkar; Jelena Diakonikolas; |
| 271 | Guided By The Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We prove that, with moderate over-parameterization, the student network undergoes a feature learning phase, where the router’s learning process are “guided by the experts, that recovers the teacher’s parameters. |
Fangshuo Liao; Anastasios Kyrillidis; |
| 272 | Robust Federated Clustering Under Heterogeneity and Adversaries Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches often exhibit significant performance degradation under these conditions and fail to return accurate solutions. To overcome these limitations, we introduce a novel federated clustering algorithm that combines client-side differential privacy with Byzantine-robust aggregation at the server, based on a novel efficient and robust clustering procedure. |
Martín Bravo; Sebastian Dalleiger; |
| 273 | Amortized Structural Variational Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose amortized structural variational inference (ASVI), which injects structural dependencies among latent variables through neural architectures that encode local neighborhood information. |
Shitao Fan; Carlos Misael Madrid Padilla; Yun Yang; Lizhen Lin; |
| 274 | Rate Optimal Learning of Equilibria from Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For the interactive setting, we introduce a framework that combines reward-free reinforcement learning with interactive MAIL and instantiate it with an algorithm, \emph{\ours}. |
Till Freihaut; Luca Viano; Emanuele Nevali; Volkan Cevher; Matthieu Geist; Giorgia Ramponi; |
| 275 | Training Latent Diffusion Models with Interacting Particle Algorithms Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel particle-based algorithm for end-to-end training of latent diffusion models. |
Tim Y. J. Wang; Juan Kuntz; O. Deniz Akyildiz; |
| 276 | Faster Parallel MCMC: Metropolis Adjustment Is Best Served Warm Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, no practical scheme uses this strategy, due to the difficulty of automatically selecting the step size during the unadjusted phase. We here develop Late Adjusted Parallel Sampler (LAPS), which is precisely such a scheme and is applicable out of the box. |
Jakob Robnik; Uros Seljak; |
| 277 | Learning Under Moral Hazard with Instrumental Regression and Generalized Method of Moments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study the foundational multitasking principal–agent contract design problem and demonstrate how instrumental regression and the generalized method of moments (GMM) estimator can be used to estimate or learn a good contract. |
Shiliang Zuo; |
| 278 | Variational Inference Via Radial Transport Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in many cases of practical interest, Gaussian distributions might not capture the correct radial profile of $\pi$, resulting in poor coverage. In this work, we approach the VI problem from the perspective of optimizing over these radial profiles. |
Luca Ghafourpour; Sinho Chewi; Alessio Figalli; Aram-Alexandre Pooladian; |
| 279 | Causal-DRF: Conditional Kernel Treatment Effect Estimation Using Distributional Random Forest Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by causal forests for CATE estimation, we develop a forest-based method to estimate the conditional kernel treatment effect (CKTE), based on the recently introduced Distributional Random Forest (DRF) algorithm. |
Jeffrey Näf; Junhyung Park; Herbert Susmann; |
| 280 | Root Cause Analysis of Outliers in Unknown Cyclic Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study the propagation of outliers in cyclic causal graphs with linear structural equations, tracing them back to one or several root cause nodes. |
Daniela Schkoda; Dominik Janzing; |
| 281 | E-Scores for (In)Correctness Assessment of Generative Model Outputs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, since these methods are based on p-values, they are susceptible to p-hacking, i.e., choosing the tolerance level post-hoc can invalidate the guarantees. We therefore leverage e-values to complement generative model outputs with e-scores as measures of incorrectness. |
Guneet S. Dhillon; Javier Gonzalez; Teodora Pandeva; Alicia Curth; |
| 282 | On The Hardness of Reinforcement Learning with Transition Lookahead Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of $\ell$ actions before deciding its course of action. |
Corentin Pla; Hugo Richard; Marc Abeille; Nadav Merlis; Vianney Perchet; |
| 283 | On The Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Mamba’s empirical performance has matched or surpassed SoTA transformers on such diverse benchmarks, the theoretical foundations underlying its powerful representational capabilities remain less explored. In this work, we investigate the expressivity of selective state-space layers using multivariate polynomials, and prove that they surpass linear transformers in expressiveness. |
Edo Cohen-Karlik; Itamar Zimerman; Liane Galanti; Ido Andrew Atad; Amir Globerson; Lior Wolf; |
| 284 | From Transformers to State Spaces: GeoMamba-SE(3) for Fast and Accurate Molecular Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To accelerate the computation and relieve the burdensome pre-training, we propose a Mamba-based framework that leverages selective state space models to learn molecular representations more efficiently. |
Jiayu Qin; Zhengquan Luo; Jian Chen; Xuhui Li; Jiayi Chen; zhiqiang xu; |
| 285 | FedCCA: Federated Canonical Correlation Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Canonical Correlation Analysis (CCA) is a key tool for cross-modal learning, but centralized solutions are impractical due to the heavy cost of high-dimensional covariance operations and the privacy sensitivity of distributed data. To address these challenges, we propose FedCCA, a federated framework that replaces explicit inverses and inner least-squares solves with a truncated von Neumann series, reducing matrix inversions to lightweight matrix–vector multiplications while retaining provable convergence. |
Zhengquan Luo; Kai Fong Ernest Chong; Pengfei Wei; Changyou Chen; Peilin Zhao; Renmin Han; Chunlai Zhou; Yunlong Wang; zhiqiang xu; |
| 286 | DIVERSED: Relaxed Speculative Decoding Via Dynamic Ensemble Verification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This constraint leads to the rejection of many plausible tokens, lowering the acceptance rate and limiting overall time speedup. To overcome this limitation, we propose DynamIc VErification RElaxed SpEculative Decoding (DIVERSED), a relaxed verification framework that improves time efficiency while preserving generation quality. |
Ziyi Wang; Siva Rajesh Kasa; Ankith M S; Santhosh Kumar Kasa; Jiaru Zou; Sumit Negi; Ruqi Zhang; Nan Jiang; Qifan Song; |
| 287 | ZipMoE: A Theoretically-Grounded Mixture of Experts Approach ForParameter-Efficient Deep Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The relentless growth of large language models (LLMs) presents formidable challenges for their training and deployment. To address this critical bottleneck, we introduce ZipMoE, a novel family of parameter-efficient building blocks inspired by the Mixture of Experts (MoE) paradigm. |
Lin Chen; Kyriakos Axiotis; Gang Fu; Kaiyuan Wang; Mohammadhossein Bateni; Vahab Mirrokni; |
| 288 | A Consequentialist Critique of Binary Classification Evaluation: Theory, Practice, and Tools Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, our empirical review of practices at major ML venues (ICML, FAccT, CHIL) reveals a dominant reliance on accuracy and AUC-ROC. To address this disconnect, we introduce a decision-theoretic framework mapping evaluation metrics to their appropriate use cases, along with a practical Python package, \texttt{briertools}, designed to make proper scoring rules more usable in real-world settings. |
Gerardo Flores; Alyssa Hasegawa Smith; Abigail E. Schiff; Julia Fukuyama; Ashia C. Wilson; |
| 289 | Deformed Decomposition for Non-negative Tensors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We presently reformulate tensor decompositions using deformed algebra, which is associated with a generalized product such that the exponential law holds for generalized exponential functions, and show that the best rank-$1$ approximation thereby reduces to a convex optimization problem for the rich $\chi$-divergence family. Building on this foundation, we propose the deformed many-body approximation for non-negative tensors, which expands model capacity while maintaining global optimality by preserving the flatness of the model manifold. |
Kazu Ghalamkari; Petr Taborsky; Morten Mørup; |
| 290 | Undersmoothing Black-Box Models for Functional Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In the classical nonparametric regression setting, we extend \texttt{Rep} with a Lepski-style method that adapts to unknown structural features of the regression function. |
Yue Yu; Debarghya Mukherjee; Moulinath Banerjee; Yaacov Ritov; |
| 291 | Adversary-Free Counterfactual Prediction Via Information-Regularized Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study counterfactual prediction under assignment bias and propose a mathematically grounded, information-theoretic approach that removes treatment–covariate dependence without adversarial training. |
Shiqin Tang; Rong Feng; Shuxin Zhuang; Youzhi Zhang; Hongzong LI; |
| 292 | Active Learning for Stochastic Contextual Linear Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an algorithm that learns a near-optimal policy by strategically sampling rewards of context-action pairs. |
Emma Brunskill; Ishani Karmarkar; Zhaoqi Li; |
| 293 | Calibrated Principal Component Regression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a new method for statistical inference in generalized linear models. |
Yixuan Wu; Yilun Zhu; Lei Cao; Naichen Shi; |
| 294 | Learning Right Monotone Permutation Matrices for Neural Subsequence Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce a neural framework that casts subsequence matching as end-to-end alignment with permutation matrices satisfying monotonicity used as differentiable approximate subsequence selectors. |
Bhavya Kohli; Soutrik Sarangi; Aziz Shameem; Abir De; |
| 295 | Optimal Learning in Games Under Delayed Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Learning in games is a central topic in both learning theory and game theory, and learning dynamics based on online learning have made significant theoretical and practical advances in recent years. |
Ruotong Zhuang; Taira Tsuchiya; Shinji Ito; |
| 296 | Understanding SAM’s Robustness to Noisy Labels Through Gradient Down-weighting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we provide a complementary explanation by analyzing SAM at the element-wise level. |
Hoang-Chau Luong; Thuc Nguyen-Quang; Dat Ba Tran; Minh-Triet Tran; |
| 297 | Towards Characterizing The Complexity of Riemannian Online Convex Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Online Convex Optimization (OCO) over Riemannian manifolds raises fundamental questions about how geometry affects algorithmic performance. While Riemannian Online Gradient … |
Hibiki Fukushima; Hiroshi Hirai; Shinji Ito; |
| 298 | Rate-optimal Design for Anytime Best Arm Identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on the framework, we propose Almost Tracking, a closed-form algorithm that has a provable guarantee on the popular risk measure. |
Junpei Komiyama; Kyoungseok Jang; Junya Honda; |
| 299 | Uncertainty Quantification for Named Entity Recognition Via Conformal Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a conformal prediction framework for NER that produces prediction sets over full label sequences with finite-sample coverage guarantees, serving an analogous role to confidence intervals in classical statistics. |
Matthew Singer; Karl Pazdernik; Srijan Sengupta; |
| 300 | Support Basis: Fast Attention Beyond Bounded Entries Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce **support-basis decomposition**, a new technique for accurate and efficient attention inference and training **without** the bounded-entry assumption. |
Maryam Aliakbarpour; Vladimir Braverman; Junze Yin; Haochen Zhang; |
| 301 | FairSHAP: Preprocessing for Fairness Through Attribution-Based Data Augmentation Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: We introduce FairSHAP, a novel preprocessing framework that leverages Shapley value attribution to improve both individual and group fairness. |
Lin Zhu; Yijun Bian; Lei You; |
| 302 | Adaptive Candidate Point Thompson Sampling for High-Dimensional Bayesian Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we introduce Adaptive Candidate Thompson Sampling (ACTS), which generates candidate points in subspaces guided by the gradient of a surrogate model sample. |
Donney Fan; Geoff Pleiss; |
| 303 | Tensor Gaussian Processes: Efficient Solvers for Nonlinear PDEs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Gaussian process (GP)/kernel-based solvers, while mathematical principled, suffer from scalability issues when handling large numbers of collocation points often needed for challenging or higher-dimensional PDEs. To overcome these limitations, we propose TGPS, a tensor-GP-based solver that introduces factor functions along each input dimension using one-dimensional GPs and combines them via tensor decomposition to approximate the full solution. |
Qiwei Yuan; Zhitong Xu; Yinghao Chen; Yiming Xu; Houman Owhadi; Shandian Zhe; |
| 304 | Implicit Updates for Average-Reward Temporal Difference Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce average-reward implicit TD($\lambda$), which employs an implicit fixed point update to provide data-adaptive stabilization while preserving the per iteration computational complexity of standard average-reward TD($\lambda$). |
Hwanwoo Kim; Dongkyu Derek Cho; Eric Laber; |
| 305 | From Restless to Contextual: A Thresholding Bandit Reformulation for Finite-horizon Improvement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Thus motivated, we introduce a reformulation of online RBs as a *budgeted thresholding contextual bandit*, which simplifies the learning problem by encoding long-term state transitions into a scalar reward. |
Jiamin Xu; Ivan Nazarov; Aditya Rastogi; Africa Perianez Santiago; Kyra Gan; |
| 306 | Beyond Pooling: Matching for Robust Generalization Under Data Heterogeneity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a matching framework that selects samples relative to an adaptive centroid and iteratively refines the representation distribution. |
Ayush Roy; Rudrasis Chakraborty; Lav R. Varshney; Vishnu Suresh Lokhande; |
| 307 | Meta-probabilistic Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we consider settings involving collections of related datasets and propose meta-probabilistic modeling (MPM) to learn the generative model structure itself. |
Kevin Zhang; Yixin Wang; |
| 308 | ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ProxRouter, which applies an exponentially tilted aggregation mechanism to balance bias and variance in nonparametric routers, improving their robustness to outliers. |
Shivam Patel; Neharika Jali; Ankur Mallick; Gauri Joshi; |
| 309 | Convexified Message-Passing Graph Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present **Convexified Message-Passing Graph Neural Networks** (CGNNs), a novel and general framework that combines the power of message-passing GNNs with the tractability of *convex* optimization. |
Saar Cohen; Noa Agmon; Uri Shaham; |
| 310 | Conformal Robust Control of Linear Systems Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: Current approaches of specifying such robust control subproblems, however, rely on hand specification of perturbations anticipated to be present upon deployment or margin methods that ignore problem structure, resulting in a lack of theoretical guarantees and overly conservative empirical performance. We, instead, propose a novel methodology for LQR systems that leverages conformal prediction to specify such uncertainty regions in a data-driven fashion. |
Yash Patel; Sahana Rayan; Ambuj Tewari; |
| 311 | Towards Blackwell Optimality: Bellman Optimality Is All You Can Get Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Further incorporating measures of immediate losses leads to the hierarchy of bias optimalities, all the way up to Blackwell optimality. In this paper, we investigate the problem of identifying policies of such optimality orders. |
Victor Boone; Adrienne Tuynman; |
| 312 | Functional Properties of The Focal-Entropy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we adopt a distributional viewpoint and study the focal-entropy, a focal-loss analogue of the cross-entropy. |
Jaimin Shah; Martina Cardone; Alex Dytso; |
| 313 | A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing formulations have two limitations: (i) observation noise, since the NTK-GP assumes noiseless targets, leading to misspecification on noisy data; (ii) the equivalence does not extend to arbitrary prior means, which are essential for well-specified models. To address (i), we introduce a regularizer into the training objective, showing its correspondence to incorporating observation noise in the NTK-GP. |
Sergio Calvo Ordoñez; Jonathan Plenk; Richard Bergna; Alvaro Cartea; José Miguel Hernández-Lobato; Konstantina Palla; Kamil Ciosek; |
| 314 | A Bayesian Information-Theoretic Approach to Data Attribution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Training Data Attribution (TDA) seeks to trace model predictions back to influential training examples, enhancing interpretability and safety. We formulate TDA as a Bayesian information-theoretic problem: subsets are scored by the information loss they induce—the entropy increase at a query when removed. |
Dharmesh Tailor; Nicolò Felicioni; Kamil Ciosek; |
| 315 | Minimax-Optimal Two-Sample Test with Sliced Wasserstein Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study the problem of nonparametric two-sample testing using the sliced Wasserstein (SW) distance. |
Binh Thuan Tran; Nicolas Schreuder; |
| 316 | On The Latent Information Geometry of The Grassmann Manifold Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we investigate the latent information geometry of deep generative models that output linear subspaces. |
Lorenzo Cazzella; Søren Hauberg; Georgios Arvanitidis; Matteo Matteucci; |
| 317 | Low Rank Based Subspace Inference for The Laplace Approximation of Bayesian Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Subspace inference for neural networks assumes that a subspace of their parameter space suffices to produce a reliable uncertainty quantification. In this work, we underpin the validity of this assumption by using low rank techniques. |
Josua Faller; Jörg Martin; |
| 318 | Spectral Text Fusion: A Frequency-Aware Approach to Multimodal Time-Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose SpecTF, a simple yet effective framework that integrates the effect of textual data on time series in the frequency domain. |
Huu Hiep Nguyen; Minh Hoang Nguyen; Dung Nguyen; Hung Le; |
| 319 | Computationally Lightweight Classifiers with Frequentist Bounds on Predictions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing kernel-based classifiers that provide such bounds scale with $\mathcal O (n^{\sim3})$ in time, making them computationally intractable for large datasets. To address this, we propose a novel, computationally efficient classification algorithm based on the Nadaraya-Watson estimator, for whose estimates we derive frequentist uncertainty intervals. |
Shreeram Murali; Cristian R. Rojas; Dominik Baumann; |
| 320 | Random Features for Operator-Valued Kernels: Bridging Kernel Methods and Neural Operators Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we investigate the generalization properties of random feature methods. |
Mike Nguyen; Nicole Mücke; |
| 321 | Power Transform Revisited: Numerically Stable, and Federated Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we find that direct implementations of power transforms suffer from severe numerical instabilities, which can lead to incorrect results or even crashes. In this paper, we provide a comprehensive analysis of the sources of these instabilities and propose effective remedies. |
Xuefeng Xu; Graham Cormode; |
| 322 | Multi-Agent Lipschitz Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our objective is to design a communication-free policy that maximizes collective reward, while separating coordination costs from learning costs. |
Sourav Chakraborty; Amit Kiran Rege; Claire Monteleoni; Lijun Chen; |
| 323 | Leveraging Machine-Learned Advice in Strategic Interactions with No-Regret Learners Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study how an agent in a two-player repeated game can effectively utilize potentially imperfect advice when interacting with a no-regret learner (i.e., satisfying a no-external or no-swap regret condition). |
Tinashe Handina; Tongxin Li; Kishan Panaganti; Eric Mazumdar; Adam Wierman; |
| 324 | Meta Sparse Principal Component Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study the meta-learning for support recovery (i.e., non-zero coordinates of the eigenvectors) in high-dimensional Principal Component Analysis. |
Imon Banerjee; Jean Honorio; |
| 325 | A Correlation Analysis Approach to Finding Interpretable Latent Representations Via Conditional Generative Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Supervised disentanglement, that is, learning interpretable nonlinear latent representations of a target data view informed by an auxiliary data view, is a central challenge in interpretable machine learning. We formulate this problem as a partially linear invertible canonical correlation analysis (PLiCCA). |
James Buenfil; Eardi Lila; |
| 326 | Scalable Policy Maximization Under Network Interference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a scalable Thompson sampling algorithm that maximizes cumulative rewards on a $n$-node network while allowing for both nodes and edges to be sampled at each time period. |
Aidan Gleich; Eric Laber; Alexander Volfovsky; |
| 327 | Fast Private Adaptive Query Answering for Large Data Domains Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce new techniques to integrate residual queries into state-of-the-art adaptive mechanisms such as AIM. |
Miguel Fuentes; Brett Mullins; Yingtai Xiao; Daniel Kifer; Cameron N Musco; Daniel Sheldon; |
| 328 | Exact and Approximate MCMC for Doubly-intractable Probabilistic Graphical Models Leveraging The Underlying Independence Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a method that does not require perfect or sequential sampling, and can be applied to both classes of methods: exact and approximate MCMC. |
Yujie Chen; Antik Chakraborty; Anindya Bhadra; |
| 329 | Optimal Local Convergence Rates of Stochastic First-Order Methods Under Local Alpha-PL Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study the local oracle complexity of stochastic first-order methods under a local $\alpha$–Polyak–Łojasiewicz ($\alpha$–PŁ) condition in a neighborhood of a target connected component $\mathcal M’$ of the local minimizer set. |
Saeed Masiha; Saber Salehkaleybar; Niao He; Negar Kiyavash; Patrick Thiran; |
| 330 | Moonwalk: Inverse-Forward Differentiation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For non-submersive layers, we introduce fragmental gradient checkpointing, which records only the minimal subset of residuals necessary to restore the cotangents erased by the Jacobian. |
Dmitrii Krylov; Armin Karamzade; Roy Fox; |
| 331 | Enforcing Fair Predicted Scores on Intervals of Percentiles By Difference-of-Convex Constraints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a novel framework for building partially fair machine learning models, which enforce fairness within a specific score range of interest, such as the middle range where decisions are most contested, while maintaining flexibility in other regions. |
Yutian He; Yankun Huang; Yao Yao; Qihang Lin; |
| 332 | Welfare-Centric Clustering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we model group utilities based on both distances and proportional representation and formalize two optimization objectives based on welfare-centric clustering: the Rawlsian (Egalitarian) objective and the Utilitarian objective. |
Claire Jie Zhang; Seyed A. Esmaeili; Jamie Heather Morgenstern; |
| 333 | A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a Thompson Sampling (TS) approach to Bayesian optimization with preferential feedback that models comparisons using a monotone link on latent utility differences and leverages the dueling kernel induced by a base kernel. |
Joseph Lazzaro; Davide Buffelli; Da-shan Shiu; Sattar Vakili; |
| 334 | Bounds and Identification of Joint Probabilities of Potential Outcomes and Observed Variables Under Monotonicity Assumptions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose new families of monotonicity assumptions and formulate the bounding problem as a linear programming problem. |
Naoya Hashimoto; Yuta Kawakami; Jin Tian; |
| 335 | OEUVRE: OnlinE Unbiased Variance-Reduced Loss Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce OEUVRE, an estimator that evaluates each incoming sample on the function learned at the current and previous time steps, recursively updating the loss estimate in constant time and memory. |
Kanad Shrikar Pardeshi; Bryan Wilder; Aarti Singh; |
| 336 | Scalable Utility-Aware Multiclass Calibration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study scalable \emph{evaluation} of multiclass calibration. |
Mahmoud Hegazy; Michael I. Jordan; Aymeric Dieuleveut; |
| 337 | Tight Analysis of Decentralized SGD: A Markov Chain Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel analysis of the Decentralized Stochastic Gradient Descent (DSGD) algorithm with constant step size, interpreting the iterates of the algorithm as a Markov chain. |
Lucas Versini; Paul Mangold; Aymeric Dieuleveut; |
| 338 | Provably Efficient Reinforcement Learning for Sparse Dynamical Systems with Non-Gaussian Noise Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study online RL in environments where the system dynamics, modeled as $s’=f(s,a)+$ noise, is assumed to be sparse with respect to a big feature map, a structural idea inspired by the SINDy framework. |
Davide Maran; Gianmarco Tedeschi; Enea Gusmeroli; Marcello Restelli; |
| 339 | Data Distribution Valuation Using Generalized Bayesian Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate the data distribution valuation problem, which aims to quantify the values of data distributions from their samples. |
Cuong N. Nguyen; Cuong V. Nguyen; |
| 340 | Causal Partial Identification Via Conditional Optimal Transport Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we demonstrate continuity of the COT problem under a stronger topology induced by the adapted Wasserstein distance. |
Sirui Lin; Zijun Gao; Jose Blanchet; Peter Glynn; |
| 341 | DP-SPRT: Differentially Private Sequential Probability Ratio Tests Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose DP-SPRT, a wrapper that can be calibrated to achieve desired error probabilities and privacy constraints, addressing a significant gap in previous work. |
Thomas Michel; Debabrota Basu; Emilie Kaufmann; |
| 342 | Certifying Reading Comprehension in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a fundamentally different approach: rather than evaluating LLMs with fixed datasets, we introduce the first framework for certifying LLMs based on large probability distributions over realistic reading comprehension prompts. |
Isha Chaudhary; Vedaant V Jain; Gagandeep Singh; |
| 343 | Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Offline reinforcement learning (RL) agents often fail when deployed, as the gap between training datasets and real environments leads to unsafe behavior. To address this, we present SAS (Self-Alignment for Safety), a transformer-based framework that enables test-time adaptation in offline safe RL without retraining. |
Seungyub Han; Hyung Jjn Kim; Jungwoo Lee; |
| 344 | A Polynomial-Time Approximation for Pairwise Fair $k$-Median Clustering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study pairwise fair $k$-Median with $\ell \ge 2$ groups, where for every cluster $C$ and every group $i \in [\ell]$, the number of points in $C$ from group $i$ must be at most $t$ times the number of points in $C$ from any other group $j \in [\ell]$, for a given integer $t$. |
Sayan Bandyapadhyay; Eden Chlamtáč; Zachary Friggstad; Mahya Jamshidian; Yury Makarychev; Ali Vakilian; |
| 345 | Fast and Robust Simulation-Based Inference With Optimization Monte Carlo Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a new method for differentiable simulators that delivers accurate posterior inference with substantially reduced runtimes. |
Vasileios Gkolemis; Christos Diou; Michael U. Gutmann; |
| 346 | Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To that end, we focus on the finite-horizon linear MDPs and propose an optimistic actor-critic framework that uses parametric log-linear policies. |
Max Qiushi Lin; Reza Asad; Kevin Tan; Haque Ishfaq; Csaba Szepesvari; Sharan Vaswani; |
| 347 | CoreSPECT: Enhancing Clustering Algorithms Via An Interplay of Density and Geometry Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we provide a novel perspective on the underlying structure of real-world data with ground-truth clustering via characterization of an abundantly observed yet often overlooked *density–geometry* correlation. |
Chandra Sekhar Mukherjee; Joonyoung Bae; Jiapeng Zhang; |
| 348 | Hellinger Multimodal Variational Autoencoders Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we revisit multimodal inference through the lens of probabilistic opinion pooling, an optimization-based approach. |
Huyen Thuc Khanh Vo; Isabel Valera; |
| 349 | On The Role of Depth in The Expressivity of RNNs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We extend our analysis to 2RNNs, a generalization of RNNs with multiplicative interactions between inputs and hidden states. |
Maude Lizaire; Michael Rizvi-Martel; Éric Dupuis; Guillaume Rabusseau; |
| 350 | On The Convergence and Straightness of Rectified Flow Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Building on this theory, we establish the first theoretical framework to analyze the straightness of RF. |
Vansh Bansal; Saptarshi Roy; Alessandro Rinaldo; Purnamrita Sarkar; |
| 351 | KQ-SVD: Compressing The KV Cache with Provable Guarantees on Attention Fidelity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior compression methods typically apply low-rank decomposition to keys alone or attempt to jointly embed queries and keys, but both approaches neglect that attention fundamentally depends on their inner products. In this work, we prove that such strategies are sub-optimal for approximating the attention matrix. |
Damien Lesens; Beheshteh T. Rakhshan; Guillaume Rabusseau; |
| 352 | Tractable Shapley Values and Interactions Via Tensor Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show how to replace the $O(2^n)$ coalition enumeration over $n$ features behind Shapley values and Shapley-style interaction indices with a \emph{few-evaluation} scheme on a tensor-network (TN) surrogate: TN-SHAP. |
Farzaneh Heidari; Chao Li; Guillaume Rabusseau; |
| 353 | ADOPT: Additive Optimal Transport Regression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel framework for additive optimal transport regression, which incorporates additive structure through optimal geodesic transports. |
Wookyeong Song; Hans-Georg Müller; |
| 354 | Three-operator Splitting with Stale Gradients for Faster Non-linear Optimal Transport Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Scalable optimization for non-linear optimal transport (OT) poses unique challenges; it requires efficient memory management of large matrices, effective parallelization strategies suited for modern accelerators like GPUs, and theoretical guarantees that support practical implementation patterns. To address these challenges, we introduce a new algorithm based on three-operator splitting that reduces gradient computation costs by allowing gradient evaluations to run asynchronously and in parallel with other computations. |
Jacob Lindbäck; David Alvarez-Melis; Mikael Johansson; |
| 355 | Generalization Bounds for Spectral GNNs Via Fourier Domain Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Spectral graph neural networks learn graph filters, but their behavior with increasing depth and polynomial order is not well understood. We analyze these models in the graph Fourier domain, where each layer becomes an element-wise frequency update, separating the fixed spectrum from trainable parameters and making depth and order explicit. |
Vahan A. Martirosyan; Daniele Malitesta; Hugues Talbot; Jhony H. Giraldo; Fragkiskos D. Malliaros; |
| 356 | Connectome-Guided Optimization for Deep Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By contrast, artificial neural networks typically rely on manually-tuned learning-rate schedules or generic adaptive optimizers whose hyperparameters remain largely agnostic to a model’s internal dynamics. In this paper, we propose Connectome-Guided Automatic Learning Rate (CG-ALR) that dynamically constructs a functional connectome of the neural network from neuron co-activations at each training iteration and adjusts learning rates online as this connectome reconfigures. |
Peilin He; Tananun Songdechakraiwut; |
| 357 | Efficient and Accurate Tensor Compression Via Recursive Sketching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose improved sketching algorithms that provide unbiased estimates for pairwise inner products, with significantly lower variance – independent of the number of modes—compared to that of Rakhshan and Rabusseau (AISTAT, 2020). |
Amit Sharma; Mohammad Azhar Khan; Rameshwar Pratap; |
| 358 | Boltzmann Exploration for Heavy-Tailed Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose heavy Boltzmann exploration (H-BE), a Boltzmann-style randomized policy whose action-selection probabilities remain available in closed form under heavy-tailed noise. |
Hyeon-jun Park; Yoon-Sik Cho; Kyungjae Lee; |
| 359 | A Semi-Supervised Kernel Two-Sample Test Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, incorporating covariates potentially breaks the exchangeability assumption under the null, which further complicates a calibration procedure. To address these issues, we propose a semi-supervised method that produces a test statistic with asymptotic normality, while effectively integrating additional information from covariates. |
Gyumin Lee; Shubhanshu Shekhar; Ilmun Kim; |
| 360 | Open Multi-agent Multi-armed Bandit with Applications in Permissionless Blockchain Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study a multi-agent multi-armed bandit problem (MA-MAB) in open systems, where multiple agents can enter and leave at any time and face multiple bandit problems to minimize the group-wise cumulative regret. |
Mengfan Xu; Diego Klabjan; |
| 361 | Minimax Generalized Cross-Entropy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a minimax formulation of generalized cross-entropy (MGCE) that results in a convex optimization over classification margins. |
Kartheek Bondugula; Santiago Mazuelas; Aritz Pérez; Anqi Liu; |
| 362 | Improving Coverage in Combined Prediction Sets with Weighted P-values Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a framework for the *weighted aggregation of prediction sets*, where weights are assigned to each prediction set based on their contribution. Our framework offers flexible control over how the sets are aggregated, achieving tighter coverage bounds that interpolate between the $1-2\alpha$ guarantee of the combined models and the $1-\alpha$ guarantee of an individual model depending on the distribution of weights. |
Gina Wong; Drew Prinster; Suchi Saria; Rama Chellappa; Anqi Liu; |
| 363 | Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we provide a rigorous analysis of the singular value spectrum of the attention matrix and establish the first Gaussian equivalence result for attention. |
Tomohiro Hayase; Benoit Collins; Ryo Karakida; |
| 364 | An Information-Geometric Approach to Artificial Curiosity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This framework provides important constraints to the engineering of intrinsic reward while integrating foundational exploration methods into a single, cohesive model. |
Alexander Nedergaard; Pablo A. Morales; |
| 365 | High-Performance Self-Supervised Learning By Joint Training of Flow Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Their iterative sampling also incurs substantial computational and energy costs, hindering industrial and edge AI applications. To address these issues, we propose the Flow Matching-based Sensor Foundation Model (SenFlow), which jointly trains a representation encoder and a conditional flow matching generator. |
Kosuke Ukita; Tsuyoshi Okita; |
| 366 | GiVA: Gradient-Informed Bases for Vector-Based Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work introduces GiVA, a gradient-based initialization strategy for vector-based adaptation. |
Neeraj Gangwar; Rishabh Deshmukh; Michael Shavlovsky; Hancao Li; Vivek Mittal; Lexing Ying; Nickvash Kani; |
| 367 | Patch2Loc: Learning to Localize Patches for Unsupervised Brain Lesion Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While supervised learning methods require annotated lesions, we propose a new unsupervised approach (Patch2Loc) that learns from normal patches taken from structural MRI. We train a neural network model to map a patch back to its spatial location within a slice of the brain volume. |
Hassan Baker; Austin J. Brockmeier; |
| 368 | Standard Acquisition Is Sufficient for Asynchronous Bayesian Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods posit that standard acquisitions lead to redundant and repeated queries, proposing complex solutions to enforce diversity in queries. |
Ben Riegler; James A C Odgers; Vincent Fortuin; |
| 369 | Parameter-Free Dynamic Regret for Unconstrained Linear Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We provide a simple approach to combining the guarantees of several bandit algorithms, allowing us to optimally adapt to the number of switches $S_T = \sum_t\mathbb{I}\\{\boldsymbol{u}\_t \neq \boldsymbol{u}\_{t-1}\\}$ of an arbitrary comparator sequence. |
Alberto Rumi; Andrew Jacobsen; Nicolò Cesa-Bianchi; Fabio Vitale; |
| 370 | Optimized Projection-Free Algorithms for Online Learning: Construction and Worst-Case Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work studies and develops projection-free algorithms for online learning with linear optimization oracles (a.k.a.~Frank–Wolfe) for handling the constraint set, and for convex loss functions. |
Julien Weibel; Pierre Gaillard; Wouter M Koolen; Adrien Taylor; |
| 371 | On Global Convergence Rates for Federated Softmax Policy Gradient Under Heterogeneous Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We provide global convergence rates for vanilla and entropy-regularized federated softmax stochastic policy gradient ($\texttt{FedPG}$) with local training. |
Safwan Labbi; Paul Mangold; Daniil Tiapkin; Eric Moulines; |
| 372 | Multiple Invertible and Partial-Equivariant Function for Latent Vector Transformation to Enhance Disentanglement in VAEs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel method, called \textit{Multiple Invertible and Partial-Equivariant Transformation} (MIPE-Transformation), which integrates two main parts: (1) \textit{Invertible and Partial-Equivariant Transformation} (IPE-Transformation), guaranteeing an invertible latent-to–transformed-latent mapping while preserving partial input-to-latent equivariance in the transformed latent space; and (2) \textit{Exponential-Family Conversion} (EF-Conversion) to extend the standard Gaussian prior to an approximate exponential family via a learnable conversion. |
Hee-Jun Jung; Jaehyoung Jeong; Kangil Kim; |
| 373 | A New Perspective on Minimum-Norm Interpolation Under Gaussian Covariates Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a new perspective on MNI under isotropic Gaussian covariates by leveraging tools from high-dimensional geometry. |
Gil Kur; Zong Shang; Paul Simanjuntak; Guillaume Lecué; Reese Pathak; |
| 374 | LLM-as-a-Judge on A Budget Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a principled variance-adaptive approach leveraging multi-armed bandit theory and concentration inequalities. |
Aadirupa Saha; Aniket Wagde; Branislav Kveton; |
| 375 | Differentially Private Linear Regression and Synthetic Data Generation with Statistical Guarantees Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a method for linear regression with valid inference under Gaussian DP. |
Shurong Lin; Aleksandra Slavkovic; Deekshith Reddy Bhoomireddy; |
| 376 | WSBD: Freezing-Based Optimizer for Quantum Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The training of Quantum Neural Networks (QNNs) is hindered by the high computational cost of gradient estimation and the barren plateau problem, where optimization landscapes become intractably flat. To address these challenges, we introduce Weighted Stochastic Block Descent (WSBD), a novel optimizer with a dynamic, parameter-wise freezing strategy. |
Christopher Kverne; Mayur Akewar; Yuqian Huo; Tirthak Patel; Janki Bhimani; |
| 377 | Adaptive Diffusion Guidance Via Stochastic Optimal Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current approaches to CFG scheduling—determining the appropriate guidance weight—are largely heuristic and lack a solid theoretical foundation. This work addresses these limitations on two fronts. |
Iskander Azangulov; Peter Potaptchik; Qinyu Li; Eddie Aamari; George Deligiannidis; Judith Rousseau; |
| 378 | On Propagation of Chaos for The Fisher-Rao Gradient Flow in Entropic Mean-Field Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such problems can be solved by constructing continuous-time gradient flows that converge to the minimizer of the energy function under consideration, and then implementing discrete-time algorithms that approximate the flow. In this work, we focus on the Fisher-Rao gradient flow and we construct an interacting particle system that approximates the flow as its mean-field limit. |
Petra Lazić; Linshan Liu; Mateusz B. Majka; |
| 379 | Finite-Time Analysis of Gradient Descent for Shallow Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we analyze a shallow Transformer with $m$ independent heads trained by projected gradient descent in the kernel regime. |
Enes Arda; Semih Cayci; Atilla Eryilmaz; |
| 380 | Atlas-based Manifold Representations for Interpretable Riemannian Machine Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we aim to give a proof of concept of the effectiveness and potential of atlas-based methods. |
Ryan Allen Robinett; Sophia Madejski; Kyle Ruark; Samantha J. Riesenfeld; Lorenzo Orecchia; |
| 381 | Parameter-Efficient Multi-Task Learning Via Progressive Task-Specific Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While these methods perform well in single-task learning, extending them to multi-task learning exacerbates common issues, such as task interference and negative transfer, due to the limited number of trainable parameters. To address these challenges, we introduce progressive task-specific multi-task adaptation, a novel parameter-efficient approach for multi-task learning. |
Neeraj Gangwar; Anshuka Rangi; Rishabh Deshmukh; Holakou Rahmanian; Yesh Dattatreya; Nickvash Kani; |
| 382 | Interpreting and Controlling Model Behavior Via Constitutions for Atomic Concept Edits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a black-box interpretability framework that learns a verifiable constitution: a natural language summary of how changes to a prompt affect a model’s specific behavior, such as its alignment, correctness, or adherence to constraints. |
Neha Kalibhat; Zi Wang; Prasoon Bajpai; Drew Proud; Wenjun Zeng; Been Kim; Mani Malek; |
| 383 | Free Random Projection for In-Context Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: We introduce Free Random Projection, an input mapping grounded in free probability theory that constructs random orthogonal matrices where hierarchical structure arises inherently. |
Tomohiro Hayase; Benoit Collins; Nakamasa Inoue; |
| 384 | Unmixing Mean Embeddings for Domain Adaptation with Target Label Proportion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel approach to domain adaptation within the context of Learning from Label Proportions (LLP). |
Alain Rakotomamonjy; Maxime Berar; Mokhtar Z. Alaya; |
| 385 | Representative, Informative, and De-Amplifying: Requirements for Robust Bayesian Active Learning Under Model Misspecification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our first contribution is to provide a mathematical analysis of generalization error in the presence of model misspecification, revealing that, beyond covariate shift, generalization error is also driven by a previously unidentified phenomenon we term *error (de-)amplification*. |
Roubing Tang; Sabina J. Sloman; Samuel Kaski; |
| 386 | FlowPINNs: A Variational Framework for PDE Parameter Inference and Uncertainty Quantification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, providing principled uncertainty quantification (UQ) for the predictions obtained using PINNs remains a significant challenge. To address this limitation, we introduce flowPINNs, a probabilistic framework for estimation and UQ in PDE parameter inverse problems. |
David Dalton; Hao Gao; Dirk Husmeier; |
| 387 | Partial Monotonicity for Submodular Maximization with A Knapsack Constraint Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we analytically show that many previously proposed algorithms for monotone submodular maximization with a knapsack constraint can achieve improved approximation guarantees under partial monotonicity with a simple modification: enforcing positive marginal gain. |
Tong Cheng; Xueyan Tang; |
| 388 | Where The Score Lives: A Wavelet View of Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A variety of architectures including CNNs, U-Nets, and Transformers have been used as the score-approximation network in such diffusion modeling; however, to date, relatively little is known about how these architectural choices impact generative behavior. In this work, to provide insight into this area, we propose an analytically solvable parameterization of the score function using an expansion in a 2D orthogonal wavelet basis. |
Emma Lucia Byrnes Finn; Binxu Wang; T. Anderson Keller; Demba E. Ba; |
| 389 | Thompson Sampling-like Algorithms for Stochastic Rising Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The strong regularity of the expected rewards in the SRRB setting suggests that specific instances may be tackled effectively using adapted and sliding-window TS approaches. This work provides novel regret analyses for such algorithms in SRRBs, highlighting the challenges and providing new technical tools of independent interest. |
Marco Fiandri; Alberto Maria Metelli; Francesco Trovò; |
| 390 | Sequential Off-Policy Learning with Logarithmic Smoothing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present and study a simple algorithm for *sequential off-policy learning*, combining Logarithmic Smoothing (LS) estimation with online PAC-Bayesian tools. |
Maxime Haddouche; Otmane Sakhi; |
| 391 | Private Synthetic Graph Generation and Fused Gromov-Wasserstein Distance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, instead of starting from a network, we start with the complex data set itself and construct both a network representation and a corresponding synthetic network generator. We build a network model directly based on the underlying complex system data, capturing its structure and attributes. |
Leoni Carla Wirth; Gholamali Aminian; Gesine Reinert; |
| 392 | Provable Target Sample Complexity Improvements As Pre‑Trained Models Scale Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we provide a theoretical investigation by introducing a novel framework, caulking, inspired by parameter-efficient fine-tuning (PEFT) methods such as adapter-based fine-tuning, low-rank adaptation, and partial fine-tuning. |
Kazuto Fukuchi; Ryuichiro Hataya; Kota Matsui; |
| 393 | Efficient Subgroup Analysis Via Optimal Trees with Global Parameter Fusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Nevertheless, these approaches encounter significant limitations, including suboptimal partitions induced by greedy heuristics and overfitting from locally estimated splits, especially under limited sample sizes. To address these limitations, we propose a fused optimal causal tree method that leverages mixed-integer optimization (MIO) to facilitate precise subgroup identification. |
Zhongming Xie; Joseph Giorgio; Jingshen Wang; |
| 394 | Sparse Linear Bandits with Fixed Sparsity Support: Adversarial and Stochastic Regimes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we design algorithms for the $l_\infty$- and $l_2$-balls that integrate sparsity support identification with the OSMD algorithm, achieving regret bounds $O(s\sqrt{T}\log T )$ and $O(\sqrt{sT}\log T )$, respectively. |
Kyoungseok Jang; Nam Phuong Tran; Nicolò Cesa-Bianchi; |
| 395 | From Guess2Graph: When and How Can Unreliable Experts Safely Boost Causal Discovery in Finite Samples? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While integrating expert knowledge (including from LLMs) as constraints promises to improve performance, guarantees for existing methods require perfect predictions or uncertainty estimates, making them unreliable for practical use. We propose the Guess2Graph (G2G) framework, which uses expert guesses to guide the sequence of statistical tests rather than replacing them. |
Sujai Hiremath; Dominik Janzing; Philipp Michael Faller; Patrick Blöbaum; Elke Kirschbaum; Shiva Kasiviswanathan; Kyra Gan; |
| 396 | Differentially Private Clustering in Data Streams Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we provide the first differentially private (DP) algorithms for $k$-means and $k$-median clustering of $d$-dimensional Euclidean data points over a stream of length at most $T$, using space that is sublinear in $T$, in the continual release setting where the algorithm is required to output a clustering at every timestep. |
Alessandro Epasto; Tamalika Mukherjee; Peilin Zhong; |
| 397 | Revisiting Social Welfare in Bandits: UCB Is (Nearly) All You Need Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We demonstrate that an initial uniform exploration phase followed by a standard Upper Confidence Bound (UCB) algorithm achieves near-optimal Nash regret, while relying only on additive Hoeffding bounds, and naturally extending to sub-Gaussian rewards. |
Dhruv Sarkar; Nishant Pandey; Sayak Ray Chowdhury; |
| 398 | High-dimensional Level Set Estimation with Trust Regions and Double Acquisition Functions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose TRLSE, an algorithm for high-dimensional LSE, which identifies and refines regions near the threshold boundary with dual acquisition functions operating at both global and local levels. |
Giang Ngo; Dat Phan Trong; Dang Nguyen; Sunil Gupta; |
| 399 | High-dimensional Learning with Noisy Labels Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on our derivations, we design an optimized method that is shown to be provably more efficient in handling noisy labels in high dimensions. |
Aymane El Firdoussi; Mohamed El Amine Seddik; |
| 400 | Topological Alignment of Shared Vision-Language Embedding Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent multilingual extensions have alleviated this gap but enforce instance-level alignment while neglecting the global geometry of the shared embedding space. We address this problem by introducing **ToMCLIP** (**To**pological Alignment for **M**ultilingual **CLIP**), a topology-aware framework aligning embedding spaces with topology-preserving constraints. |
Junwon You; Kang Dasol; Jae-Hun Jung; |
| 401 | Narrowing Action Choices with AI Improves Human Sequential Decisions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Can we use the same principle to achieve complementarity in sequential decision making tasks? In this paper, we answer this question affirmatively. |
Eleni Straitouri; Stratis Tsirtsis; Ander Artola Velasco; Manuel Gomez Rodriguez; |
| 402 | The Riemannian Geometry Associated to Gradient Flows of Linear Convolutional Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study geometric properties of the gradient flow for learning deep linear convolutional networks. |
El Mehdi Achour; Kathlén Kohn; Holger Rauhut; |
| 403 | Optimal Posterior Sampling for Policy Identification in Tabular Markov Decision Processes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a randomized and computationally efficient algorithm for best policy identification that combines posterior sampling with an online learning algorithm to guide exploration in the MDP. |
Cyrille Kone; Kevin Jamieson; |
| 404 | Graphon Mixtures Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a generative model that captures both hub and dense structures. |
Sevvandi Kandanaarachchi; Cheng Soon Ong; |
| 405 | Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency Without Model Sweeps Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our approach introduces Voronoi-type loss functions aligned with the gate-partition geometry and establishes finite-sample convergence rates for the maximum likelihood estimator (MLE). |
TienHai Do; Trung Nguyen Mai; TrungTin Nguyen; Nhat Ho; Binh T. Nguyen; Christopher Drovandi; |
| 406 | On The Intrinsic Dimensions of Data in Kernel Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The second, denoted $d_K$, is the effective dimension, derived from the decay rate of Kolmogorov $n$-widths associated with $K$ on $\Omega$. Given a probability measure $\mu$ on $\Omega$, we analyze the relationship between these $n$-widths and eigenvalues of the integral operator $\phi \mapsto \int_\Omega K(\cdot,x)\phi(x)\,d\mu(x)$. |
Rustem Takhanov; |
| 407 | LAMP: Extracting Local Decision Surfaces From Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce **LAMP** (**L**ocal **A**ttribution **M**apping **P**robe), a method that shines light onto a black-box language model’s decision surface and studies how reliably a model maps its stated reasons to its reported predictions by approximating a decision surface. |
Ryan Chen; Youngmin Ko; Catherine Cho; Zeyu Zhang; Mauro Giuffrè; Sunny Chung; Dennis Shung; Bradly C. Stadie; |
| 408 | Longitudinal Flow Matching for Trajectory Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose \textit{Interpolative Multi-Marginal Flow Matching} (IMMFM), a framework that learns continuous stochastic dynamics jointly consistent with multiple observed time points. |
Mohammad Mohaiminul Islam; Thijs P. Kuipers; Sharvaree Vadgama; Coen de Vente; Afsana Khan; Clara I. Sánchez; Erik J Bekkers; |
| 409 | Laplace Approximation for Bayesian Variable Selection Via Le Cam’s One-step Procedure Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While many existing methods offer strong statistical guarantees, they are often computationally intractable in high-dimensional problems. To address this issue, we introduce a novel Laplace approximation method based on Le Cam’s one-step procedure, termed \textsf{OLAP}. |
Tianrui Hou; Yves Atchade; |
| 410 | Learning to Bid in Discriminatory Auctions with Budget Constraints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Exploiting a decomposition of utility across units, we develop polynomial-time learning algorithms based on shortest paths in a directed acyclic graph, obtaining sublinear regret under both full-information and bandit feedback. |
Negin Golrezaei; Sourav Sahoo; |
| 411 | A Modularized Framework for Piecewise-Stationary Restless Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the resulting exploration–detection delay trade-off, we propose a modular framework that integrates arbitrary RMAB base algorithms with change detection and a novel diminishing exploration mechanism. |
Kuan-Ta Li; Chia-Chun Lin; Ping-Chun Hsieh; Yu-Chih Huang; |
| 412 | Generalized and Optimal Straight-Through Estimators Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a principled axiomatic approach to define a general family of gradient estimators and show that it subsumes many existing methods. |
James Hooper; Alexander Shekhovtsov; |
| 413 | Pure Exploration with Infinite Answers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study pure exploration problems where the set of correct answers is possibly infinite, e.g., the regression of any continuous function of the means of the bandit. |
Riccardo Poiani; Martino Bernasconi; Andrea Celli; |
| 414 | A Projection-based Framework for Gradient-free and Parallel Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a feasibility-seeking approach to neural network training. |
Andreas Bergmeister; Manish Krishan Lal; Stefanie Jegelka; Suvrit Sra; |
| 415 | Learning Hyperparameters Via A Data-Emphasized Variational Objective Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: In this paper, we study gradient-based learning of hyperparameters via the evidence lower bound (ELBO) objective from Bayesian variational methods. |
Ethan Harvey; Mikhail Petrov; Michael C Hughes; |
| 416 | PAC-Bayesian Bounds on Constrained $f$-Entropic Risk Measures Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: PAC generalization bounds on the risk, when expressed in terms of the expected loss, are often insufficient to capture imbalances between subgroups in the data. To tackle this limitation, we introduce a new family of risk measures, called constrained $f$-entropic risk measures, which enable finer control over distributional shifts and subgroup imbalances via $f$-divergences, and include the Conditional Value at Risk (CVaR), a well-known risk measure. |
Hind Atbir; Farah Cherfaoui; Guillaume Metzler; Emilie Morvant; Paul Viallard; |
| 417 | Composable Coresets for Constrained Determinant Maximization and Beyond Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study algorithms for construction of *composable coresets* for the task of *Determinant Maximization* under *partition constraint*. |
Sepideh Mahabadi; Thuy-Duong Vuong; |
| 418 | Explicit Density Approximation for Neural Implicit Samplers Using A Bernstein-Based Convex Divergence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose \emph{dual-ISL}, obtained by interchanging the roles of the target $p$ and model density $\tilde p$ within ISL, which induces a \emph{convex} optimization problem over model densities. |
José Manuel de Frutos; Pablo M. Olmos; Manuel A. Vázquez; Joaquin Miguez; |
| 419 | RealStats: A Rigorous Real-Only Statistical Framework for Fake Image Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce a rigorous, statistically grounded framework for fake image detection that focuses on producing a probability score interpretable with respect to the real-image population. |
Haim Zisman; Uri Shaham; |
| 420 | In-Context Learning for Discrete Optimal Transport: Can Transformers Sort? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent progress has shed light on the mechanisms underlying in-context learning in statistical tasks: language models can implement linear regression and classification by iteratively extracting features at test time. This naturally raises a broader question: *Can we analyze ICL beyond statistical learning and extend it to discrete algorithmic tasks relevant to NLP? |
Hadi Daneshmand; |
| 421 | Empirical PAC-Bayes Bounds for Markov Chains Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we prove a new PAC-Bayes bound for Markov chains. |
Vahe Karagulyan; Pierre Alquier; |
| 422 | BOAT: Navigating The Sea of in Silico Predictors for Antibody Design Via Multi-Objective Bayesian Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present BOAT, a versatile Bayesian optimization framework for multi-property antibody engineering. |
Jackie Rao; Ferran Gonzalez; Leon Gerard; Alexandra Gessner; |
| 423 | TESLA: Taylor Expansion of Sinusoidal Learnable Activations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high-order components. |
Daehwa Ko; JaeHyeon Kim; SeungHyun Ham; Jay Hoon Jung; |
| 424 | Fast and Robust Convergence Rate for TD(0) with Linear Function Approximation, Universal Learning Steps and I.I.D. Samples Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study the finite-time behavior of the TD(0) temporal-difference method with linear function approximation (LFA). |
Ziad Kobeissi; Eloïse Berthier; |
| 425 | PowerSoftmax: Towards Secure LLM Inference Over Encrypted Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a new HE-friendly variant of self-attention that offers a stable form for training and is easy to approximate with polynomials for secure inference. |
Itamar Zimerman; Allon Adir; Ehud Aharoni; Matan Avitan; Moran Baruch; Nir Drucker; Jenny Lerner; Ramy Masalha; Reut Moshe; Omri Soceanu; |
| 426 | Empirically Calibrated Conditional Independence Tests Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Empirically Calibrated Conditional Independence Tests (ECCIT), a method that measures and corrects for miscalibration. |
Milleno Pan; Antoine de Mathelin; Wesley Tansey; |
| 427 | On The Bias of Variational Resampling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We also show that the truncation bias implies that the particle approximation of the target distribution is restricted to a region in which the unnormalised weights are larger than some threshold with high probability. We prove that this probability approaches $1$ if $M = \mathrm{O}(N)$ as $N \to \infty$. |
Axel Finke; Oskar Kviman; Nicola Branchini; Víctor Elvira; |
| 428 | TLDR: Network Inversion for Extreme-Case Training-Like Data Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the trained model weights are frequently distributed under the assumption that sharing model parameters does not compromise the confidentiality or privacy of the training data. In this work, we challenge this assumption by presenting \textbf{Training-Like Data Reconstruction (TLDR)}, as a general-purpose and architecture-agnostic framework for reconstructing training data from a fully trained classifier. |
Pirzada Suhail; Sunny Gupta; Amit Sethi; |
| 429 | Regularizing Attention Scores with Bootstrapping Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Leveraging statistical learning techniques, we introduce the bootstrapping for attention scores which generates a baseline distribution of attention scores by resampling input features. |
Neo Christopher Chung; Maxim Laletin; |
| 430 | Learning to Explore With Lagrangians For Bandits Under Unknown Constraints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Pure exploration in bandits formalises multiple real-world problems, such as tuning hyper-parameters or conducting user studies to test a set of items, where different safety, resource, and fairness constraints on the decision space naturally appear. We study these problems as pure exploration in multi-armed bandits with unknown linear constraints, where the aim is to identify an *$r$-optimal and feasible policy* as fast as possible with a given level of confidence. |
Udvas Das; Debabrota Basu; |
| 431 | UniPROT: Uniform Prototype Selection Via Partial Optimal Transport with Submodular Guarantees Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present \texttt{UniPROT}, a novel subset selection framework that minimizes the optimal transport (OT) distance between a uniformly weighted prototypical distribution and the target distribution. |
Prateek Chanda; Prayas Agrawal; Karthik S. Gurumoorthy; Ganesh Ramakrishnan; Bamdev Mishra; Pratik Jawanpuria; |
| 432 | On The Identifiability of Tensor Ranks Via Prior Predictive Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces a rigorous approach to determine rank identifiability in probabilistic tensor models, based on prior predictive moment matching. |
Eliezer de Souza da Silva; Arto Klami; Diego Mesquita; Iñigo Urteaga; |
| 433 | Numerical Fragility in Transformers: A Layer-wise Theory for Risk Estimation and Selective Stabilization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a first-order decomposition of output mismatch into layer-local attention, LayerNorm, and residual-transport terms, and derive from it a practical causal risk estimator and a budgeted controller, Bound-Guided Selective Stabilization (BGSS). |
Jinwoo Baek; |
| 434 | Bandits in Flux: Adversarial Constraints in Dynamic Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate the challenging problem of adversarial multi-armed bandits operating under time-varying constraints, a scenario motivated by numerous real-world applications. To address this complex setting, we propose a novel primal-dual algorithm that extends online mirror descent through the incorporation of suitable gradient estimators and effective constraint handling. |
Tareq Si Salem; |
| 435 | Provable FDR Control for Deep Feature Selection: Deep MLPs and Beyond Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a flexible feature selection framework based on deep neural networks that approximately controls the false discovery rate (FDR), a measure of Type-I error. |
Kazuma Sawaya; |
| 436 | On The Normalization of Confusion Matrices: Methods and Geometric Interpretations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Because confusion matrix values jointly reflect both factors, it is difficult to disentangle their individual effects. To address this issue, we introduce bi-normalization via Iterative Proportional Fitting, a generalization of row and column normalization. |
Johan Erbani; Sonia Ben Mokhtar; Pierre-Edouard Portier; Elöd Egyed-Zsigmond; Diana Nurbakova; |
| 437 | Predictive Deep Sets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Curiously, their success still primarily stems from the expressiveness of transformers, lacking a bias for modeling the functional structures between features and labels shared across datasets. We argue and show that this leads to training sample inefficiency and sub-optimal performance, and address this by introducing a novel set encoding technique called Predictive Deep Sets. |
Alex Hämäläinen; Sammie Katt; Samuel Kaski; |
| 438 | Differentially Private and Federated Structure Learning in Bayesian Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce Fed-Sparse-BNSL, a novel federated method for learning linear Gaussian Bayesian network structures that addresses both challenges. |
Ghita Fassy El Fehri; Aurélien Bellet; Philippe Bastien; |
| 439 | Deep Feedback Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This feedback mechanism introduces dynamics into otherwise static architectures, enabling DFMs to iteratively refine their internal state and mimic aspects of biological decision making. We model this process as a differential equation solved through a recurrent neural network, stabilized via exponential decay to ensure convergence. |
David Calhas; Arlindo L. Oliveira; |
| 440 | Uncovering Hidden Training Dynamics in Neural Networks Via Inter-Sample Influence Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Influence Graphs (IGs), directed inter-sample graphs where each edge weight $w_{ij}$ quantifies how optimizing on sample $X_i$ influences the loss of sample $X_j$. |
Dylan Tan Hong Tai; Jiayin Zhang; Rohan Ghosh; Mehul Motani; |
| 441 | Convex Markov Games and Beyond: New Proof of Existence, Characterization and Learning Algorithms for Nash Equilibria Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While cMGs expand the modeling frontier, their theoretical foundations, particularly the structure of Nash equilibria (NE) and guarantees for learning algorithms, are not yet well understood. In this work, we address these gaps for an extension of cMGs, which we term General Utility Markov Games (GUMGs), capturing new applications requiring coupling between agents’ occupancy measures. |
Anas Barakat; Ioannis Panageas; Antonios Varvitsiotis; |
| 442 | Accelerated Learning on Large-Scale Screens Using Generative Library Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this article, we introduce algorithms to optimize high throughput screens for data creation and model training. |
Eli N Weinstein; Andrei Slabodkin; Mattia Gollub; Elizabeth Baker Wood; |
| 443 | Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing high-probability convergence analyses typically require the clipping threshold to increase with the number of optimization steps, which is incompatible with standard DP mechanisms like the Gaussian mechanism. In this work, we close this gap by providing the first high-probability convergence analysis for DP-Clipped-SGD with a fixed clipping level, applicable to both convex and non-convex smooth optimization under heavy-tailed noise, characterized by a bounded central $\alpha$-th moment assumption, $\alpha \in (1,2]$. |
Saleh Vatan Khah; Savelii Chezhegov; Shahrokh Farahmand; Samuel Horváth; Eduard Gorbunov; |
| 444 | Tractable Gaussian Phase Retrieval with Heavy Tails and Adversarial Corruption with Near-Linear Sample Complexity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study efficient algorithms for robust phase retrieval with heavy-tailed noise when a constant fraction of both the measurements $y_i$ and the sensing vectors $a_i$ may be arbitrarily adversarially corrupted. |
SANTANU DAS; jatin batra; |
| 445 | On The Hardness of Auditing Model Properties Under Updates: Complexity of Property-Preserving Updates Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose a generic algorithmic framework for efficient PAC auditing, powered by an Empirical Property Optimization (EPO) oracle. |
Ayoub Ajarra; Debabrota Basu; |
| 446 | Two Mathematical Models of Knowledge Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Many hypotheses compete to explain the successes of knowledge distillation. To help address this, we propose and analyze a mathematical model of distillation, which suggests that distillation’s performance comes not from obtaining better models but from easier to optimize landscapes. |
Audrey Xie; Ludwig Schmidt; John Duchi; |
| 447 | Adaptive A/B Testing Under Nonstationary Dynamics Using State-Space Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Response-adaptive randomization (RAR) design provides a natural alternative, adaptively allocating participants over time based on accrued information. In this work, we propose a methodology framework that addresses these challenges. |
Junzhe Shao; Waverly Wei; Jingshen Wang; |
| 448 | Is Supervised Learning Really That Different From Unsupervised? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We demonstrate how supervised learning can be decomposed into a two-stage procedure, where (1) all model parameters are selected in an unsupervised manner, and (2) the outputs y are added to the model, without changing the parameter values. |
Oskar Allerbo; Thomas B. Schön; |
| 449 | BASTION: A Bayesian Framework for Trend and Seasonality Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce BASTION (Bayesian Adaptive Seasonality and Trend DecompositION), a flexible Bayesian framework for decomposing time series into trend and multiple seasonality components. |
Jason B. Cho; David S. Matteson; |
| 450 | CONTEXTUAL RANKING AND MATCHING. OPTIMAL REGRET UNDER LST Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Additionally, we assume that the strength of each player can be inferred from some available contextual information through the contextualised linear stochastic transitivity model \textbf{(LST)}. We propose an algorithm that performs matchmaking by selecting pairs of maximum informativeness among admissible pairs and prove that its regret is optimal up to logarithmic factors. |
Hafedh El Ferchichi; Vianney Perchet; Matthieu LERASLE; |
| 451 | Learning Markov Processes As Sum-of-Square Forms for Analytical Belief Propagation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes a functional modeling framework leveraging sparse Sum-of-Squares (SoS) forms for valid (conditional) density estimation. |
Peter Amorese; Morteza Lahijanian; |
| 452 | Monotone and Conservative Policy Iteration Beyond The Tabular Case Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Reliable Policy Iteration (RPI) and Conservative RPI (CRPI), variants of Policy Iteration (PI) and Conservative PI (CPI), that retain tabular guarantees under function approximation. |
S.R. Eshwar; Gugan Thoppe; Ananyabrata Barua; Aditya Gopalan; Gal Dalal; |
| 453 | Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, both approaches often assume uniform annotator preferences and rely on binary comparisons, overlooking two key limitations: the diversity of human evaluators and the limitations of pairwise feedback. In this work, we address both these issues. |
Keertana Chidambaram; Karthik Vinay Seetharaman; Vasilis Syrgkanis; |
| 454 | The Unseen Adversaries: Robust and Generalized Defense Against Adversarial Patches Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The current research need arises from the above vulnerabilities and the lack of efforts to tackle these two singularities independently and, especially, in combination. In this research, we have, for the first time, combined these two prominent singularities and proposed a novel dataset. |
Vishesh Kumar; Akshay Agarwal; |
| 455 | Efficient Model Performance Evaluation Using A Combination of Expert and Crowd-sourced Labels Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Expert human labelers are high-quality but scarce and resource-intensive to obtain, while crowd-sourced labels are more readily accessible at scale but lower in quality. We propose Maven (Model And Voter EvaluatioN), a hierarchical Bayesian model that combines these two label sources to produce model performance estimates on binary tasks that are less biased than using crowd-sourced labels alone and have lower variance than using expert labels alone. |
Sam Corbett-Davies; Viet-An Nguyen; Udi Weinsberg; |
| 456 | Policy Learning with Abstention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The ability to abstain, that is, to defer to a safe default or an expert, is crucial but largely unexplored in this context. To remedy this, we introduce a framework for policy learning with abstention, in which policies that choose not to assign a treatment to some customers/patients receive a small, additive reward on top of the value of a random guess. |
Ayush Sawarni; Jikai Jin; Justin Whitehouse; Vasilis Syrgkanis; |
| 457 | Latent-IMH: Efficient Bayesian Inference for Inverse Problems with Approximate Operators Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this framework, we introduce Latent-IMH, a sampling method based on the Metropolis-Hastings independence (IMH) sampler. |
Youguang Chen; George Biros; |
| 458 | Q-ShiftDP: A Differentially Private Parameter-Shift Rule for Quantum Machine Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce the Differentially Private Parameter-Shift Rule (Q-ShiftDP), the first privacy mechanism tailored to QML. |
Hoang M. Ngo; Nhat Hoang-Xuan; Quan Minh Nguyen; Nguyen Hoang Khoi Do; Incheol Shin; My T. Thai; |
| 459 | A Unifying Framework for Unsupervised Concept Extraction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present a unified theoretical framework for unsupervised concept extraction, in which we frame the task of concept extraction as identifying a generative model. |
Chandler Squires; |
| 460 | Replicable Machine Learning: Theory and Algorithms for Stochastic Convex and Non-Convex Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Replicable algorithms produce identical outputs with high probability when run on independent samples drawn from the same distribution, providing strong reproducibility guarantees for machine learning pipelines. We study replicability in machine learning in Vapnik’s general learning setting, which encompasses stochastic optimization over convex and non-convex loss classes, establishing algorithms with near-optimal sample complexity across these settings. |
Raman Arora; Kaibo Zhang; |
| 461 | Model Selection for Average Reward RL with Application to Utility Maximization in Repeated Games Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose $\textsf{MRBEAR}$, an online model selection algorithm for the average reward RL setting which is based on the idea of regret balancing and elimination. |
Alireza Masoumian; James R. Wright; |
| 462 | Regression Descent: A Statistical Framework for Neural Network Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Regression Descent (RD), a novel optimization algorithm for training deep neural networks that reformulates each gradient step as a regression problem in the span of the Jacobian. |
Kamaljeet Singh; Nicolas Hengartner; Hao Zhang; Brian Wesley Bell; James Hyman; |
| 463 | Adaptive Memory Momentum Via A Model-Based Framework for Deep Learning Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce an adaptive memory mechanism that replaces constant momentum with a dynamic momentum coefficient that is adjusted online during optimization. |
Kristi Topollai; Anna Ewa Choromanska; |
| 464 | CADENT: Gated Hybrid Distillation for Sample-Efficient Transfer in Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Context-Aware Distillation with Experience-gated Transfer (CADENT), a framework that unifies strategic automaton-based knowledge with tactical policy-level knowledge into a coherent guidance signal. |
Mahyar Alinejad; Yue Wang; George K. Atia; |
| 465 | Beyond Spectral Clustering: Probabilistic Cuts for Differentiable Graph Partitioning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a unified probabilistic framework that covers a wide class of cuts, including Normalized Cut. |
Ayoub Ghriss; |
| 466 | Value Gradient Sampler: Learning Invariant Value Functions for Equivariant Diffusion Sampling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose the Value Gradient Sampler (VGS), a diffusion sampler parameterized by value functions. |
Himchan Hwang; Hyeokju Jeong; Dong Kyu Shin; Che-Sang Park; Sehee Kweon; Sangwoong Yoon; Frank C. Park; |
| 467 | Securing Model Weights Against Eavesdropping Adversaries in Federated Learning Using Quantization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our lightweight, architecture-agnostic approach combines low-bit quantization with an adaptive clipping rule to thwart reconstruction attacks, even under warm adversary initialization. |
Kushal Chakrabarti; Dipankar Maity; |
| 468 | Confidence-Guided Self-Training for Gradual Domain Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a theoretical framework for self-training under gradual domain shift that explicitly quantifies and controls the pseudo-labeling error incurred at each round. |
Akram Heidarizadeh; Akram Awad; HanQin Cai; George K. Atia; |
| 469 | Adversarial Debiasing for Parameter Recovery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, prediction errors from machine learning models can lead to bias in downstream estimation tasks, including regression. In this paper, we show how this bias can arise, propose a test for detecting bias, and demonstrate the use of an adversarial machine learning algorithm in order to generate predictions suitable for unbiased downstream estimation. |
Luke Sanford; Megan Ayers; Matthew Gordon; Eliana Stone; |
| 470 | Black-Box Optimization From Small Offline Datasets Via Meta Learning with Synthetic Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes {\em Surrogate Learning with Optimization Bias via Synthetic Task Generation} (\textsc{OptBias}), a meta-learning framework that directly tackles data scarcity. |
Azza Fadhel; The Hung Tran; Trong Nghia Hoang; Jana Doppa; |
| 471 | The Information Geometry of Local Generalization Dynamics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Information-theoretic bounds on generalization are foundational to learning theory, yet their static form offers limited insight into the dynamic, iterative nature of modern optimization. |
Emmanouil M. Athanasakos; |
| 472 | Robustness and Generalization in Uncertainty-Aware Message Passing Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We address a more realistic scenario where noise or finite measurement precision introduces uncertainties in node feature values. |
Alesia Chernikova; Moritz Laber; Narayan G. Sabhahit; Tina Eliassi-Rad; |
| 473 | Frequency-Based Hyperparameter Selection in Games Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a principled approach to hyperparameter selection in games by leveraging frequency estimation of oscillatory dynamics. |
Aniket Sanyal; Baraah A. M. Sidahmed; Rebekka Burkholz; Tatjana Chavdarova; |
| 474 | Optimal Transport Guarantees to Nonparametric Regression for Locally Stationary Time Series Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a conditional probability distribution estimator for LSTS through Nadaraya–Watson (NW) kernel smoothing. |
Jan Nino G. Tinio; Mokhtar Z. Alaya; Salim Bouzebda; |
| 475 | Clustering-Based Edge Augmentation for Minimizing The Kirchhoff Index Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study the problem of augmenting a given graph by adding $k$ edges to minimize the Kirchhoff index. |
Prasanth Yalamanchili; Aditya Bhaskara; |
| 476 | High-Probability Bounds for Heterogeneous Local Differential Privacy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Departing from the common in-expectation analyses, and for one-dimensional and multi-dimensional mean estimation problems, we develop finite sample upper bounds in $\ell_2$-norm that hold with probability at least $1-\beta$. |
Maryam Aliakbarpour; Alireza Fallah; Swaha Roy; Ria Stevens; |
| 477 | It’s All In The (Exponential) Family: An Equivalence Between Maximum Likelihood Estimation and Control Variates For Sketching Algorithms Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Maximum likelihood estimators (MLE) and control variate estimators (CVE) have been used in conjunction with known information across sketching algorithms and applications in machine learning. |
Keegan Kang; Kerong Wang; Ding Zhang; Rameshwar Pratap; Bhisham Dev Verma; Benedict H. W. Wong; |
| 478 | Optimistic Reinforcement Learning with Quantile Objectives Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we develop UCB-QRL, an optimistic learning algorithm for the $\tau$-quantile objective in finite-horizon Markov decision processes (MDPs). |
Mohammad Alipour-Vaezi; Huaiyang Zhong; Kwok-leung Tsui; Sajad Khodadadian; |
| 479 | Recovery Guarantees for Continual Learning of Dependent Tasks: Memory, Data-Dependent Regularization, and Data-Dependent Weights Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: At the heart of developing CL theory lies the challenge that the data distribution varies across tasks, and we argue that properly addressing this challenge requires understanding this variation–dependency among tasks. To explicitly model task dependency, we consider nonlinear regression tasks and propose the assumption that these tasks are dependent in such a way that the data of the current task is a nonlinear transformation of previous data. |
Liangzu Peng; Uday Kiran Reddy Tadipatri; Ziqing Xu; Eric Eaton; Rene Vidal; |
| 480 | Refining Covariance Matrix Estimation in Stochastic Gradient Descent Through Bias Reduction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While classical methods—such as plug-in and batch-means estimators—are available, they either require inaccessible second-order (Hessian) information or suffer from slow convergence. To address these challenges, we propose a novel, fully online de-biased covariance estimator that eliminates the need for second-order derivatives while significantly improving estimation accuracy. |
Ziyang Wei; Wanrong Zhu; Jingyang Lyu; Wei Biao Wu; |
| 481 | Scalable Model-Based Clustering with Sequential Monte Carlo Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel SMC algorithm that decomposes clustering problems into approximately independent subproblems, allowing a more compact representation of the algorithm state. |
Connie Trojan; Pavel Myshkov; Paul Fearnhead; James Hensman; Tom Minka; Christopher Nemeth; |
| 482 | Identification and Estimation of Probabilities of Causation in The Presence of Confounding and Selection Bias Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Thus, they are not applicable to studies in the presence of confounding and selection biases. To address this problem, this paper provides novel identification conditions for the probabilities of causation by using (i) two proxy covariates and (ii) an instrumental variable and a proxy covariate. |
Ryusei Shingaki; Haruka Yoshida; manabu kuroki; |
| 483 | T$_k$CP: Context-Aware Pooling Via Top-k% Activation Selection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This limits performance on tasks requiring both fine localization and holistic understanding. To address this, we propose Top-$k$\% Contextual Pooling (TkCP), a framework that preserves informative activations based on contextual importance. |
Seo-Yeon Choi; Kyungsu Lee; |
| 484 | Loss Gaps Parity for Fairness in Heterogeneous Federated Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose EAGLE, a novel federated learning algorithm that explicitly regularizes the global model to minimize disparities in loss gaps across clients. |
Brahim Erraji; Michaël Perrot; Aurélien Bellet; |
| 485 | Efficient Inference for Coupled Hidden Markov Models in Continuous Time and Discrete Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Latent Interacting Particle Systems, a model class parameterizing the generator of each Markov chain in the system. |
Giosue Migliorini; Padhraic Smyth; |
| 486 | Adaptive Combinatorial Experimental Design: Pareto Optimality for Decision-Making and Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we provide the first investigation into adaptive combinatorial experimental design, focusing on the trade-off between regret minimization and statistical power in combinatorial multi-armed bandits (CMAB). |
Xie Hongrui; Junyu Cao; Kan Xu; |
| 487 | Learnability with Partial Labels and Adaptive Nearest Neighbors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we mathematically characterize the scenarios in which PLL is feasible. |
Nicolás A. Errandonea; Santiago Mazuelas; Jose A. Lozano; Sanjoy Dasgupta; |
| 488 | A Goemans-Williamson Type Algorithm for Identifying Subcohorts in Clinical Trials Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our theoretical contribution is a rounding technique, similar to that of Goemans and Wiliamson (1995), which approximates the optimal solution within a factor of $0.82$. |
Pratik Worah; |
| 489 | Set to Be Fair: Demographic Parity Constraints for Set-Valued Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we address the problem of set-valued classification under demographic parity and expected size constraints. |
Eyal Haïm Cohen; Christophe Denis; Mohamed Hebiri; |
| 490 | Fast Quasar-Convex Optimization with Constraints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Obtaining an accelerated algorithm that makes nearly optimal $\bigotilde{1/(\gamma\sqrt{\epsilon})}$ first-order queries to a $\gamma$-quasar convex smooth function \emph{with constraints} was independently asked as an open problem in Martinez-Rubio (2022); Lezane, Langer and Koolen, (2024). In this work, we solve this question by designing an inexact accelerated proximal point algorithm that we implement using a first-order method achieving the aforementioned rate and, as a consequence, we improve the complexity of the accelerated geodesically Riemannian optimization solution in Martínez-Rubio (2022). |
David Martínez-Rubio; |
| 491 | Near-optimal Rank Adaptive Inference of High Dimensional Matrices Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an algorithm that combines a Least-Squares estimator with a universal singular value thresholding procedure. |
Frédéric Zheng; Yassir Jedra; Alexandre Proutiere; |
| 492 | Lower Bounds for Public-Private Learning Under Distribution Shift Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work we extend the known lower bounds for public-private learning to a setting where the two data sources exhibit significant distribution shift. |
Amrith Setlur; Pratiksha Thaker; Jonathan Ullman; |
| 493 | Stick-Breaking Embedded Topic Model with Continuous Optimal Transport for Online Analysis of Document Streams Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although these methods naturally align with real-world scenarios, they have received considerably less attention from the community compared to their offline counterparts, due to specific additional challenges. To tackle these issues, we present SB-SETM, an innovative model extending the Embedded Topic Model (ETM) to process data streams by merging models formed on successive partial document batches. |
Federica Granese; Serena Villata; Charles Bouveyron; |
| 494 | Mask-Conditional Conformal Prediction: Valid Uncertainty For All Missing Data Mechanisms Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we adapt split CP to handle missing values by proposing a preimpute-mask-then-correct framework that can offer valid coverage. |
JIARONG FAN; Juhyun Park; Thi Phuong Thuy Vo; Nicolas J-B. Brunel; |
| 495 | Ergodic and Subhomogeneous Dynamics in Hyperbolic Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We analyze the long term behavior of hyperbolic neural networks through subhomogeneous layer maps, focusing on stability, growth control, and robustness under stochastic perturbations. |
Nico Alvarado; Sebastian Burgos; |
| 496 | Structured Matching Via Cost-Regularized Unbalanced Optimal Transport Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Choosing such a cost is particularly challenging when datasets live in heterogeneous spaces, often motivating practitioners to adopt Gromov–Wasserstein formulations. To address this challenge, we introduce cost-regularized unbalanced optimal transport (CR-UOT), a framework that allows the ground cost to vary while allowing mass creation and removal. |
Emanuele Pardini; Katerina Papagiannouli; |
| 497 | Minimizing Human Intervention in Online Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: When the horizon $T$ is at least exponential in the embedding dimension $d$, the geometry of the class regions can be learned. In this regime, we propose the Conservative Hull-based Classifier (CHC), which maintains convex hulls of expert-labeled queries and calls the expert when a query lands outside all known hulls. |
William Réveillard; Vasileios Saketos; Alexandre Proutiere; Richard Combes; |
| 498 | Multilayer Correlation Clustering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We establish Multilayer Correlation Clustering, a novel generalization of Correlation Clustering to the multilayer setting. |
Atsushi Miyauchi; Florian Adriaens; Francesco Bonchi; Nikolaj Tatti; |
| 499 | Rank Lifting and Random Non-Linear Maps Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Deep neural networks exhibit improved training and generalization performance as the number of parameters grows well beyond the size of the training set, contradicting classical intuitions about overfitting. In order to gain a better understanding of this “benign overparameterization”, we analyze the representational capacity of a random one-hidden-layer perceptron with Gaussian weights, no bias and threshold activations. |
Andrea Drago; Maria Sofia Bucarelli; Francesco Caso; Marius Michetti; Federico Siciliano; Fabrizio Silvestri; Luca Becchetti; |
| 500 | Gradient Regularized Natural Gradients Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Gradient-Regularized Natural Gradients (GRNG), a family of scalable second-order optimizers that integrate explicit gradient regularization with natural gradient updates. |
Satya Prakash Dash; Hossein Abdi; Wei Pan; Samuel Kaski; Mingfei Sun; |
| 501 | Momentum SVGD-EM for Accelerated Maximum Marginal Likelihood Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recently, a significant body of work has adopted this perspective, leading to interacting particle algorithms for MMLE. In this paper, we propose an accelerated version of one such procedure, based on Stein variational gradient descent (SVGD), by introducing Nesterov acceleration in both the parameter updates and in the space of probability measures. |
Adam Rozzio; Rafael Athanasiades; O. Deniz Akyildiz; |
| 502 | Counterfactual Explanations Via Latent Structure for Time Series Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose CELT, a model-agnostic CF generation method for time-series classifiers, including non-differentiable and one-class models. |
Akihiro Yamaguchi; Shizuo Kaji; Kaname Matsue; Ryusei Shingaki; |
| 503 | Lag Operator SSMs: A Geometric Framework for Structured State Space Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a direct, first-principles framework for constructing discrete-time SSMs that is both flexible and modular. |
Sutashu Tomonaga; Kenji Doya; Noboru Murata; |
| 504 | From Hawkes Processes to Attention: Time-Modulated Mechanisms for Event Sequences Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing Transformer-based methods mostly inject temporal information only via positional encodings, relying on shared or parametric decay structures, which limits their ability to capture heterogeneous and type-specific temporal effects. Inspired by this observation, we derive a novel attention operator called Hawkes Attention from the multivariate Hawkes process theory for MTPP, using learnable per-type neural kernels to modulate query, key and value projections, thereby replacing the corresponding parts in the traditional attention. |
Xinzi Tan; kejian zhang; Junhan Yu; Doudou Zhou; |
| 505 | Towards Sensitivity-Aware Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, it remains unclear how sensitivity awareness relates to established notions of privacy, such as differential privacy (DP), thereby making it difficult to deploy meaningfully in real-world applications. In this work, we formalize the notion of sensitivity awareness and theoretically establish its connection to DP. |
Dren Fazlija; Iyiola Emmanuel Olatunji; Daniel Kudenko; Sandipan Sikdar; |
| 506 | Generalized Correlation Shifting for Lasso Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The Lasso has been widely used in a high-dimensional setting, but its estimation accuracy may become inadequate when the covariates are highly correlated or when the number of covariates is extremely large. To overcome this problem, we propose a novel preconditioner that adaptively induces a low-rank structure in the design matrix. |
Izuru Miyazaki; Hironori Fujisawa; |
| 507 | From Cells to Sentences: An End-to-End Framework for Table Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These issues cause even state-of-the-art language models to fail at seemingly simple questions. We present a robust framework for table understanding that explicitly handles these challenges through three coordinated mechanisms: structure-aware encoders that learn invariance to common corruptions, trainable slots that compress evidence to a fixed-size representation, and grounding modules that align each slot to supporting text passages. |
Deepak Vijaykeerthy; Arvind Agarwal; |
| 508 | On The Optimal Regret of Collaborative Personalized Linear Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a two-stage collaborative algorithm that achieves the optimal regret. |
Bruce Huang; Ruida Zhou; Lin F. Yang; Suhas Diggavi; |
| 509 | Efficient Uncoupled Learning Dynamics with $\tilde{O}\left(T^{-1/4}\right)$ Last-Iterate Convergence in Bilinear Saddle-Point Problems Over Convex Sets Under Bandit Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study last-iterate convergence of learning algorithms in bilinear saddle-point problems, a preferable notion of convergence that captures the day-to-day behavior of learning dynamics. |
Arnab Maiti; Claire Jie Zhang; Kevin Jamieson; Jamie Heather Morgenstern; Ioannis Panageas; Lillian J. Ratliff; |
| 510 | From Token Imbalance to Balanced Routing: An ELBO-Regularized Probabilistic Framework for Contrastive Multimodal Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce CoPRIME (Contrastive Probabilistic Routing for IMbalanced tokens with ELBO-regularized mixture of experts), a probabilistic routing framework for multimodal representation learning that generalizes multimodal representation learning beyond vision-text by tackling the fundamental challenge of extreme token imbalance across modalities. |
Habibeh Naderi; Behrouz Haji Soleimani; Stan Matwin; |
| 511 | Injecting Measurement Information Yields A Fast and Noise-Robust Diffusion-Based Inverse Problem Solver Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose to estimate the conditional posterior mean $\mathbb{E} [\mathbf{x}_0 | \mathbf{x}_t, \mathbf{y}]$, which can be formulated as the solution to a lightweight, single-parameter maximum likelihood estimation problem. |
Jonathan Patsenker; Henry Li; Myeongseob Ko; Ruoxi Jia; Yuval Kluger; |
| 512 | Fundamental Limits of Non-Adaptive Group Testing With Markovian Correlation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In the sparse infections regime, where the expected number of infections scales as $O(n^{\theta})$ with $\theta \in (0,1)$, we propose a non-adaptive testing strategy with an efficient decoding algorithm. |
Aditya Narayan Ravi; Ilan Shomorony; |
| 513 | On The Misinformation in A Statistical Experiment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We demonstrate that the commonly-accepted axioms of experimental utility, such as Blackwell monotonicity, fail under misspecification, and that information measures proposed to handle it, like the Expected Generalized Information Gain (EGIG), do not obey these axioms. To resolve this, we propose a generalized axiomatic framework for robust Bayesian experimental design. |
Jake Callahan; Tommie Catanach; |
| 514 | LatticeVision: Image to Image Networks for Modeling Non-Stationary Spatial Data Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: In this work we focus on a popular class of parametric, spatially autoregressive (SAR) models. |
Antony Sikorski; Michael Ivanitskiy; Nathan Lenssen; Douglas Nychka; Daniel McKenzie; |
| 515 | Bayesian Inverse Transition Learning: Learning Dynamics from Near-Optimal Trajectories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We consider the problem of estimating the transition dynamics from near-optimal expert trajectories in the context of offline model-based reinforcement learning. |
Leo Benac; Abhishek Sharma; Sonali Parbhoo; Finale Doshi-Velez; |
| 516 | ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ConMeZO, a novel zeroth‑order optimizer that accelerates convergence by adaptive directional sampling. |
Lejs Deen Behric; Liang Zhang; Bingcong Li; Kiran Koshy Thekumparampil; |
| 517 | Information-Theoretic Error Bounds for Source Localization in Neural Sensing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formulate a point-source localization problem in $d$ dimensions, where a source inside the ball of radius $R$ emits a signal that is picked up by various sensors located at the surface of the ball. |
Leighton Pate Barnes; Yuxin Guo; Alex Dytso; Pulkit Grover; |
| 518 | Active Subspaces in Infinite Dimension Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this article, we extend this methodology to real-valued functionals on Hilbert space. |
Poorbita Kundu; Nathan Wycoff; |
| 519 | Partial VOROS: A Cost-aware Performance Metric for Binary Classifiers with Precision and Capacity Constraints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The usual area under the curve metric also does not reflect asymmetric costs for false positives and false negatives. In this paper we address all three of these issues. |
Christopher Ratigan; Kyle Heuton; Carissa Wang; Lenore Cowen; Michael C Hughes; |
| 520 | Conditional Flow Matching for Bayesian Posterior Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a generative multivariate posterior sampler via flow matching. |
Percy S. Zhai; Sowon Jeong; Veronika Rockova; |
| 521 | Reconciling Communication Compression and Byzantine-Robustness in Distributed Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce RoSDHB, a new algorithm that integrates classical Polyak momentum with a coordinated compression strategy. |
Diksha Gupta; Antonio Honsell; Chuan Xu; Nirupam Gupta; Giovanni Neglia; |
| 522 | Efficient Swap Regret Minimization in Combinatorial Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast to the weaker notion of external regret minimization — a problem which is fairly well understood in the literature — achieving no-swap regret with a polylogarithmic dependence on $N$ has remained elusive in combinatorial bandits. Our paper resolves this challenge, by introducing a no-swap-regret learning algorithm with regret that scales polylogarithmically in $N$ and is tight for the class of combinatorial bandits. |
Andreas Kontogiannis; Vasilis Pollatos; Panayotis Mertikopoulos; Ioannis Panageas; |
| 523 | AlphaFold’s Bayesian Roots in Probability Kinematics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To demonstrate this framework with precision, we introduce a tractable synthetic model in which an angular random walk prior is updated with distance-based evidence via PK, directly mirroring AlphaFold’s mechanism. |
Thomas Hamelryck; Kanti V. Mardia; |
| 524 | Reinforcement Learning Using Known Invariances Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a symmetry-aware variant of optimistic least-squares value iteration (LSVI), which leverages invariant kernels to encode invariance in both rewards and transition dynamics. |
Alexandru Cioba; Aya Kayal; Laura Toni; Sattar Vakili; Alberto Bernacchia; |
| 525 | A Divergence-Based Method for Weighting and Averaging Model Predictions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper uses a minimum divergence framework to introduce a new way of calculating model weights that can be used to average probabilistic predictions from statistical and machine learning models. |
Olav Benjamin Vassend; |
| 526 | Hypergraph Neural Networks Accelerate MUS Enumeration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes a domain-agnostic method to accelerate MUS enumeration using Hypergraph Neural Networks (HGNNs). |
Hiroya Ijima; Koichiro Yawata; |
| 527 | Asymptotic Optimality Theory of Confidence Intervals of The Mean Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We address the classical problem of constructing confidence intervals (CIs) for the mean of a distribution, given $N$ i.i.d. samples, such that the CI contains the true mean with probability at least $1 – \delta$, where $\delta \in (0,1)$. |
VIKAS DEEP; Achal Bassamboo; Sandeep Kumar Juneja; |
| 528 | Community-Enhanced Semi-seeded Network Alignment (CESSNA): A Robust Method with Application to Microbiome Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Community-Enhanced Semi-seeded Network Alignment (CESSNA), a novel algorithm that integrates community structure and partial seed information to improve matching efficiency and accuracy. |
Sijing Yu; Daniel L. Sussman; Vince Lyzinski; |
| 529 | Sharp Risk Bounds for Early-stopping in Gaussian Linear Regression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study early-stopped mirror descent (ESMD) for high-dimensional Gaussian linear regression over arbitrary convex bodies and design matrices, where the task is to minimize the in-sample mean squared error. |
Tobias Wegel; Gil Kur; Patrick Rebeschini; |
| 530 | Robust Estimation of Heterogeneous Treatment Effects in Randomized Trials Leveraging External Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Randomized trials are typically designed to detect average treatment effects but often lack the statistical power to uncover individual-level treatment effect heterogeneity, limiting their value for personalized decision-making. To address this, we propose the QR-learner, a model-agnostic learner that estimates conditional average treatment effects (CATE) within the trial population by leveraging external data from other trials or observational studies. |
Rickard K.A. Karlsson; Piersilvio De Bartolomeis; Issa Dahabreh; Jesse H. Krijthe; |
| 531 | Transportability Without Graphs: A Bayesian Approach to Identifying S-Admissible Backdoor Sets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a Bayesian method that combines observational data from the target domain with experimental data from a different domain to identify s-admissible backdoor sets, which enable unbiased estimation of causal effects across populations, without requiring the causal graph. |
Konstantina Lelova; Gregory F Cooper; Sofia Triantafillou; |
| 532 | Closed-Form Coordinate Ascent Variational Inference for Student-t Process Regression with Student-t Likelihood Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the first tractable variational inference framework for this model. |
Keisuke Onoue; Takatomi Kubo; Kazushi Ikeda; |
| 533 | DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current approaches in this domain present a clear trade-off: PPO-style methods (e.g., GRPO/DAPO) offer training stability but exhibit slow learning trajectories due to their trust-region constraints on policy updates, while REINFORCE-style approaches (e.g., CISPO) demonstrate improved learning efficiency but suffer from performance instability as they clip importance sampling weights while still permitting non-zero gradients outside the trust-region. To address these limitations, we introduce DISPO, a simple yet effective REINFORCE-style algorithm that decouples the up-clipping and down-clipping of importance sampling weights for correct and incorrect responses, yielding four controllable policy update regimes. |
Batuhan K. Karaman; Aditya Rawal; Mohammad Ghavamzadeh; Suhaila Shakiah; Arijit Biswas; Ruida Zhou; |
| 534 | Fair Clustering Via Hierarchical Fair-Dirichlet Prior Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this article, we offer a novel model-based formulation of fair clustering, complementing the existing literature which is almost exclusively based on optimizing appropriate objective functions. |
Abhisek Chakraborty; Anirban Bhattacharya; Debdeep Pati; |
| 535 | Non-Asymptotic Generalization and Optimization Bounds for Stochastic Gauss-Newton in Deep Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we analyze a stochastic Gauss-Newton (SGN) method with Levenberg-Marquardt damping and mini-batch sampling for training overparameterized deep neural networks with smooth activations in a regression setting. |
Semih Cayci; |
| 536 | Feature Importance Via Sets of Locally Performant Linear Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose $\ell\text{-MCR}$, a local feature importance method that identifies meaningful neighborhoods around a point of interest, regions where the model or data behavior is locally stable and interpretable. |
Fatemeh Tohidian; Davin Hill; Aria Masoomi; Peter J. Castaldi; Jennifer Dy; |
| 537 | Low-Complexity and Consistent Graphon Estimation from Multiple Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a new histogram-based estimator with low algorithmic complexity that achieves high accuracy by jointly aligning the nodes of all graphs, in contrast to most conventional methods that order nodes graph by graph. |
Roland Boniface Sogan; Tabea Rebafka; |
| 538 | An Information-Theoretic Approach to Understanding Transformers’ In-Context Learning of Variable-Order Markov Chains Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study transformers’ in-context learning of variable-length Markov chains (VOMCs), focusing on the finite-sample accuracy as the number of in-context examples increases. |
Ruida Zhou; Chao Tian; Suhas Diggavi; |
| 539 | Catoni-Style Change Point Detection for Regret Minimization in Piecewise-Stationary Heavy-Tailed Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we tackle the heavy-tailed piecewise-stationary bandit problem. |
Gianmarco Genalti; Sujay Bhatt; Nicola Gatti; Alberto Maria Metelli; |
| 540 | Deconfounding Scores and Representation Learning for Causal Effect Estimation with Weak Overlap Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This is especially challenging in high dimensions: the curse of dimensionality can make overlap implausible. To address this, we propose a class of feature representations called deconfounding scores, which preserve both identification and the target of estimation; the classical propensity and prognostic scores are two special cases. |
Oscar Clivio; Alexander Nicholas D’Amour; Alexander Franks; David Bruns-Smith; Christopher C. Holmes; Avi Feller; |
| 541 | Improved Algorithms for Clustering with Noisy Distance Oracles Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we design algorithms with improved guarantees for $k$-means and $k$-center clustering problems in the weak-strong oracle model. |
Pinki Pradhan; Anup Bhattacharya; Ragesh Jaiswal; |
| 542 | A Pure Hypothesis Test for Inhomogeneous Random Graph Models Based on A Kernelised Stein Discrepancy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we develop a KSD-type test for IRG models that can be carried out with a single observation of the network. |
Anum Fatima; Gesine Reinert; |
| 543 | Beyond Binary Out of Distribution Detection: Characterizing Distributional Shifts with Multi-Statistic Diffusion Trajectories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work argues that scalar-based methods are thus insufficient for OOD data to be properly contextualized and prospectively exploited, a limitation we overcome with the introduction of DISC: Diffusion-based Statistical Characterization. |
Achref Jaziri; Martin Rogmann; Martin Mundt; Visvanathan Ramesh; |
| 544 | General Weighted Averaging in Stochastic Gradient Descent: CLT and Adaptive Optimality Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Stochastic Gradient Descent (SGD) is a cornerstone of machine learning, prized for its efficiency in large-scale optimization. This paper revisits SGD by introducing a general weighted averaging framework that significantly enhances its applicability. |
Ziyang Wei; Wanrong Zhu; Wei Biao Wu; |
| 545 | Linear Convergence of The Frank-Wolfe Algorithm Over Product Polytopes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We apply our results to the problem of approximately finding a feasible point in a polytope intersection in high-dimensions, and demonstrate the practical efficiency of our algorithms through empirical results. |
Gabriele Iommazzo; David Martínez-Rubio; Francisco Criado; Elias Samuel Wirth; Sebastian Pokutta; |
| 546 | Sample Average Approximation for Alpha-Divergence Minimization with Exponential Convergence Guarantees Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To illustrate our theoretical results, we propose an implementable Sample Average Approximation algorithm that solves a discrete approximation of the original problem. |
François Bertholom; François Roueff; randal douc; |
| 547 | Examining The Bias of In-Batch Sampling in Similarity Learning with Two-Tower Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Nevertheless, in-batch sampling introduces its own issues, such as mistaking positives for negatives and oversampling popular pairs, resulting in significant performance degradation. In this work, we provide the first systematic analysis of these issues, showing that they all arise from the inconsistency between the expected objective under in-batch sampling and the full-data objective. |
Yaxu Liu; Li-Chung Lin; Chih-Jen Lin; |
| 548 | On Different Notions of Redundancy in Conditional-Independence-Based Discovery of Graphical Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Often there are tests that were not used in the construction of the graph. In this work, we show that these _redundant_ tests have the potential to _detect_ or sometimes _correct_ errors in the learned model. |
Philipp Michael Faller; Dominik Janzing; |
| 549 | RoseCDL: Robust and Scalable Convolutional Dictionary Learning for Rare-event and Anomaly Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce RoseCDL, a novel CDL algorithm designed for robust and scalable modeling of signal pattern distribution. |
Jad Yehya; Mansour Benbakoura; Cédric Allain; Benoît Malézieux; Matthieu Kowalski; Thomas Moreau; |
| 550 | Information Hidden in Gradients of Regression with Target Noise Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our key insight is a simple variance calibration: injecting Gaussian noise so that the total target noise variance equals the batch size ensures that the empirical gradient covariance closely approximates the Hessian, even when evaluated far from the optimum. |
Arash Jamshidi; Katsiaryna Haitsiukevich; Kai Puolamäki; |
| 551 | Incorporating Expert Knowledge Into Bayesian Causal Discovery of Mixtures of Directed Acyclic Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a causal elicitation strategy for heterogeneous settings, based on Bayesian experimental design (BED) principles, and a _variational mixture structure learning_ (VaMSL) method—extending the earlier _differentiable Bayesian structure learning_ (DiBS) method—to iteratively infer mixtures of causal Bayesian networks (CBNs). |
Zachris Björkman; Jorge Loria; Sophie Wharrie; Samuel Kaski; |
| 552 | Scalable Learning of Multivariate Distributions Via Coresets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, available methods are limited in their ability to handle large-scale data. We address this issue by developing a novel coreset construction for multivariate conditional transformation models (MCTMs) to enhance their scalability and training efficiency. |
Zeyu Ding; Katja Ickstadt; Nadja Klein; Alexander Munteanu; Simon Omlor; |
| 553 | Consistent PCA and Spectral Clustering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, if the output results significantly change upon the addition of new data points, it can lead to several issues such as instability in the downstream task or a lack of trust in the findings. To address these problems, we consider online variants of PCA and spectral clustering, and show that a natural subspace-preserving regularizer provides provable approximation and consistency guarantees. |
Satoshi Hara; Yuichi Yoshida; |
| 554 | Regularizing Extrapolation in Causal Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a unified framework that directly penalizes the level of extrapolation, replacing the current practice of a hard non-negativity constraint with a soft constraint and corresponding hyperparameter. |
David Arbour; Harsh Parikh; Bijan A Niknam; Elizabeth Stuart; Kara E. Rudolph; Avi Feller; |
| 555 | $k$-PCA for (non-squared) Euclidean Distances: Deterministic Polynomial Time Approximation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We provide the first polynomial-time deterministic algorithm whose both running time and approximation factor are not exponential in $k$. |
Daniel Greenhut; Dan Feldman; |
| 556 | Optimal Rates for Density and Mode Estimation with Expand-and-sparsify Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In mode estimation, we provide simple algorithms on top of our density estimator that recover single or multiple modes at optimal rates up to logarithmic factors under mild conditions. |
Kaushik Sinha; Christopher Tosh; |
| 557 | Mixture Proportion Estimation and Weakly-supervised Kernel Test for Conditional Independence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose novel assumptions based on conditional independence (CI) given the class label, which ensure identifiability even when irreducibility does not hold. |
Yushi Hirose; Akito Narahara; Takafumi Kanamori; |
| 558 | Generalizing Behavior Via Inverse Reinforcement Learning with Closed-Form Reward Centroids Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel, principled criterion that selects the average policy among those induced by the rewards in a certain bounded subset of the feasible set. |
Filippo Lazzati; Alberto Maria Metelli; |
| 559 | Learning in Continuous State-Space MDPs for Network Inventory Management Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our primary contribution is an integrated framework establishing and leveraging the Lipschitz property of the long-run average cost function. |
Hansheng Jiang; Shunan Jiang; Zuo-Jun Shen; |
| 560 | GeoTTER: Leveraging Local Geometry of Optimal Transport for Zero-Shot Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present GeoTTER, a novel framework that redefines optimal transport in the realm of zero-shot classification. |
Wei-Yang Alex Lee; Rudrasis Chakraborty; Vishnu Suresh Lokhande; |
| 561 | Prior Shift Estimation for Positive Unlabeled Data Through The Lens of Kernel Embedding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel direct estimator of the class prior which avoids estimation of posterior probabilities in both populations and has a simple geometric interpretation. |
Jan Mielniczuk; Wojciech Rejchel; Paweł Teisseyre; |
| 562 | Harnessing The Power of Reinforcement Learning for Adaptive MCMC Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Instead, we propose a novel reward based on the contrastive divergence, whose superior performance in the context of RLMH is demonstrated. |
Congye Wang; Matthew A Fisher; Heishiro Kanagawa; Wilson Ye Chen; Chris J. Oates; |
| 563 | Bandit-based Maximum Inner Product Search with Data-Dependent Confidence Intervals Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a data-dependent bandit-based algorithm in which the lengths of the confidence intervals are adaptively adjusted based on observed samples. |
Yoichi Sasaki; Yuzuru Okajima; |
| 564 | Optimal Variance and Covariance Estimation Under Differential Privacy in The Add-Remove Model and Beyond Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study the problem of estimating the variance and covariance of datasets under differential privacy in the add-remove model. |
Shokichi Takakura; Seng Pei Liew; Satoshi Hasegawa; |
| 565 | FedDuA: Doubly Adaptive Federated Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we formalize the central server optimization procedure through the lens of mirror descent and propose a novel framework, called FedDuA, which adaptively selects the global learning rate based on both inter-client and coordinate-wise heterogeneity in the local updates. |
Shokichi Takakura; Seng Pei Liew; Satoshi Hasegawa; |
| 566 | Policy Testing in Markov Decision Processes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We first derive an instance-dependent lower bound that any reasonable algorithm must satisfy, characterized as the solution to an optimization problem with non-convex constraints. Guided by this formulation, we propose a new algorithm. |
Kaito Ariu; Po-An Wang; Alexandre Proutiere; Kenshi Abe; |
| 567 | Neural Variance-aware Dueling Bandits with Deep Representation and Shallow Exploration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the first variance-aware algorithms for contextual dueling bandits that leverage shallow exploration strategies with neural networks for nonlinear utility approximation. |
Youngmin Oh; Jinje Park; Taejin Paik; |
| 568 | Provable Guarantees for Estimating Covariances Between Latent Variables with Application to Precision Matrix Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we address a fundamental challenge: How can one estimate the covariance between variables that are not directly observable? |
Haichi Long; Qifan Song; Jean Honorio; |
| 569 | Robust Estimation of A Sparse Linear Model: Provable Guarantees with Non-convexity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we address the problem of sparse regression vector estimation in the presence of corrupted samples, with a particular focus on accurately identifying the support. |
Deepak Maurya; Adarsh Barik; Jean Honorio; |
| 570 | Differentially Private Minimum Spanning Tree in Euclidean Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study benchmark problems in privately analyzing geometric graphs obtained from high dimensional embeddings. |
Zongrui Zou; Alessandro Epasto; Chenglin Fan; Rudrajit Das; |
| 571 | DeepRV: Accelerating Spatiotemporal Inference with Pre-trained Neural Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce DeepRV, a neural-network surrogate that replaces GP prior sampling, while closely matching full GP accuracy at inference including hyperparameter estimates, and reducing computational complexity to $\mathcal{O}(N^2)$, increasing scalability and inference speed. |
Jhonathan Navott; Daniel Jenson; Seth Flaxman; Elizaveta Semenova; |
| 572 | Secretary Problem with Predictions and Ordering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces a novel framework for the secretary problem that leverages predictions for both valuation and strategic scheduling. |
Kang Yiming; |
| 573 | Dualformer: Time-Frequency Dual Domain Learning for Long-term Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This issue arises due to undifferentiated propagation of frequency components across layers, causing a progressive attenuation of high-frequency information crucial for capturing fine-grained temporal variations. To address this limitation, we propose Dualformer, a principled dual-domain framework that rethinks frequency modeling from a layer-wise perspective. |
Jingjing Bai; Yoshinobu Kawahara; |
| 574 | Variance Reduction Methods Do Not Need to Compute Full Gradients: Improved Efficiency Through Shuffling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we explore popular VR techniques and propose an approach that eliminates the necessity for expensive full gradient calculations. |
Daniil Medyakov; Gleb Molodtsov; Savelii Chezhegov; Alexey Rebrikov; Aleksandr Beznosikov; |
| 575 | Orientability of Causal Relations in Time Series Using Summary Causal Graphs and Faithful Distributions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present conditions that guarantee the orientability of micro-level edges between temporal variables given the background knowledge encoded in a summary causal graph and assuming having access to a faithful and causally sufficient distribution with respect to the true unknown graph. |
Timothée Loranchet; Charles K. Assaad; |
| 576 | Preference-based Conditional Treatment Effects and Policy Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a new preference-based framework for conditional treatment effect estimation and policy learning, built on the Conditional Preference-based Treatment Effect (CPTE). |
Dovid Parnas; Mathieu Even; Julie Josse; Uri Shalit; |
| 577 | Prior Knowledge Makes It Possible: From Sublinear Graph Algorithms to LLM Test-Time Methods Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, it is not clear how much pre-training knowledge is required to answer queries with a small number of augmentation steps, which is a desirable property in practice. To address this question, we formulate multi-step reasoning as an $s$-$t$ connectivity problem on a knowledge graph. |
Avrim Blum; Daniel Hsu; Cyrus Rashtchian; Donya Saless; |
| 578 | Where You Place The Norm Matters: From Prejudiced to Neutral Initializations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We provide a theoretical characterization of how normalization choice and placement (Pre-Norm vs. Post-Norm) determine the distribution of class predictions at initialization, ranging from unbiased (Neutral) to highly concentrated (Prejudiced) regimes. |
Emanuele Francazi; Francesco Pinto; Aurelien Lucchi; Marco Baity-Jesi; |
| 579 | Happiness As A Measure of Fairness Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel fairness framework grounded in the concept of _happiness_, a measure of the utility each group gains from decision outcomes. |
Georg Pichler; Marco Romanelli; Pablo Piantanida; |
| 580 | Multi-Component VAE with Gaussian Markov Random Field Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the Gaussian Markov Random Field Multi-Component Variational AutoEncoder, embedding Gaussian Markov Random Fields into both prior and posterior distributions to explicitly model cross-component relationships. |
Fouad Oubari; Mohamed El Baha; Raphaël Meunier; Rodrigue Décatoire; Mathilde MOUGEOT; |
| 581 | On The Weight Density of L2-Regularized Linear Classification and Regression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we rigorously prove that for $L_2$-regularized support vector classification/regression, the theoretical optimum can indeed be sparse when the data have sparse feature values. |
He-Zhe Lin; Zhi-Bao Lu; Sheng-Wei Chen; Cheng-Hung Liu; Chih-Jen Lin; |
| 582 | Design-Based Finite-Sample Analysis for Regression Adjustment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a design-based, non-asymptotic framework for analyzing the regression-adjusted ATE estimator under complete randomization. |
Dogyoon Song; |
| 583 | TAS-EGNN: Task-Aware Spectral Ego-Graphs for Efficient GNNs-Based Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a Task-Aware Spectral Ego-Graph Neural Network (TAS-EGNN) that scores nodes within lightweight ego-graphs by combining (i) local spectral complexity, (ii) predictive uncertainty, and (iii) supervised error signals, followed by a greedy coverage step to avoid redundancy. |
Mebarka Allaoui; Rachid Hedjam; Sonia Gupta; |
| 584 | Unsupervised Ensemble Learning Through Deep Energy-based Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paradigm is crucial in scenarios where evaluating individual classifier performance or understanding their strengths is challenging due to limited information. We propose a novel deep energy-based method for constructing an accurate meta-learner using only the predictions of individual learners, potentially capable of capturing complex dependence structures between them. |
Ariel Maymon; Yanir Buznah; Uri Shaham; |
| 585 | Counterfactual Credit Guided Bayesian Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce Counterfactual Credit Guided Bayesian Optimization (CCGBO), a novel framework that explicitly quantifies the contribution of individual historical observations through counterfactual credit. |
Qiyu Wei; Haowei Wang; Richard Allmendinger; Mauricio A Álvarez; |
| 586 | Probabilistic Multi-dimensional Classification with Incomplete Data at The Prediction Time Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present theoretical results leading to a new probabilistic approach with efficient learning and prediction algorithms that address scalability and robustness issues. |
Thu Ha DO; Vu-Linh Nguyen; Yves Grandvalet; |
| 587 | Partially Lazy Gradient Descent for Smoothed Online Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce \textsc{$k$-lazyGD}, an online learning algorithm that bridges the gap between greedy Online Gradient Descent (OGD, for $k=1$) and lazy GD/dual-averaging (for $k=T$), creating a spectrum between reactive and stable updates. |
Naram Mhaisen; George Iosifidis; |