惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

爱范儿
爱范儿
B
Blog RSS Feed
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
量子位
博客园 - 三生石上(FineUI控件)
博客园 - 【当耐特】
Attack and Defense Labs
Attack and Defense Labs
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
人人都是产品经理
人人都是产品经理
酷 壳 – CoolShell
酷 壳 – CoolShell
Apple Machine Learning Research
Apple Machine Learning Research
阮一峰的网络日志
阮一峰的网络日志
大猫的无限游戏
大猫的无限游戏
T
Tailwind CSS Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
罗磊的独立博客
V
Visual Studio Blog
博客园 - Franky
博客园 - 叶小钗
有赞技术团队
有赞技术团队
IT之家
IT之家
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园_首页
J
Java Code Geeks
S
SegmentFault 最新的问题
Last Week in AI
Last Week in AI
月光博客
月光博客
博客园 - 司徒正美
小众软件
小众软件
The Cloudflare Blog
宝玉的分享
宝玉的分享
博客园 - 聂微东
WordPress大学
WordPress大学
雷峰网
雷峰网
V
V2EX
Engineering at Meta
Engineering at Meta
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
L
LangChain Blog
Jina AI
Jina AI
Hugging Face - Blog
Hugging Face - Blog
The Register - Security
The Register - Security
腾讯CDC
Microsoft Azure Blog
Microsoft Azure Blog
Recent Announcements
Recent Announcements
D
Docker
F
Fortinet All Blogs
美团技术团队
H
Help Net Security
U
Unit 42
MyScale Blog
MyScale Blog

cs.DS updates on arXiv.org

PAC Learning with Bandit Feedback: Sharp Sample Complexity in the Realizable Setting Algorithms with Polynomially-Improved Approximation Factors for the $2 \rightarrow q$ Norm, and Applications A computational phase transition for learning-to-sample from Ising models Covering vertices by sequential stars Fermi-Dirac machines as quantizations of neurons A Comprehensive Evaluation of Vertex Elimination Algorithms for Algorithmic Differentiation A Tight Bound on Localization of Electrical Flows Optimal Dimension-Free Sampling for Regularized Classification Reducing the Randomness in Partition Oracles for Bounded Degree Minor-Free Graphs Beyond the Half-Approximation: Fair and Efficient Online Class Matching Efficient Uniform Sampling of Surjections via their Profiles Tractable Maximization of Budgeted Phylogenetic Diversity on Networks Utilizing Node Scanwidth Fairness in Aggregation: Optimal Top-$k$ and Improved Full Ranking Learning-Augmented Online Scheduling with Parsimonious Preemption Entropy Equivalence Testing Lumberjack: Better Differentially Private Random Forests through Heavy Hitter Detection in Trees The Secretary Problem with a Stochastic Precursor Polynomial-Time Robust Multiclass Linear Classification under Gaussian Marginals Efficient Banzhaf-Based Data Valuation for $k$-Nearest Neighbors Classification Block-Sphere Vector Quantization An Approximation Algorithm for Graph Label Selection Iterative Chow Filtering for Learning with Distribution Shift Complexity of Non-Log-Concave Sampling in Fisher Information Stochastic Matching via Local Sparsification Finite Sample Bounds for Learning with Score Matching What is Learnable in Valiant's Theory of the Learnable? Provable Quantization with Randomized Hadamard Transform Min-Max Optimization Requires Exponentially Many Queries Fast and Compact Graph Cuts for the Boykov-Kolmogorov Algorithm A proximal gradient algorithm for composite log-concave sampling Adaptive Multi-Round Allocation with Stochastic Arrivals The tractability landscape of diffusion alignment: regularization, rewards, and computational primitives Mistake-Bounded Language Generation Positional LSH: Binary Block Matrix Approximation for Attention with Linear Biases Learning-Augmented Scalable Linear Assignment Problem Optimization via Neural Dual Warm-Starts A Note on Non-Negative $L_1$-Approximating Polynomials Curvature Beyond Positivity: Greedy Guarantees for Arbitrary Submodular Functions Convex Optimization with Nested Evolving Feasible Sets On the Complexity of the Matching Problem of Regular Expressions with Backreferences Simple KNN-Based Outlier Detection Achieves Robust Clustering Online Allocation with Unknown Shared Supply Equivalence of Coarse and Fine-Grained Models for Learning with Distribution Shift Accelerated Relax-and-Round for Concave Coverage Problems Contrastive Identification and Generation in the Limit Quantizing With Randomized Hadamard Transforms: Efficient Heuristic Now Proven Nearly Optimal Attention Coresets On Computing Total Variation Distance Between Mixtures of Product Distributions Exact and Approximate Algorithms for Polytree Learning Provable Accuracy Collapse in Embedding-Based Representations under Dimensionality Mismatch New Bounds for Kernel Sums via Fast Spherical Embeddings Unlearning Offline Stochastic Multi-Armed Bandits Matroid Algorithms Under Size-Sensitive Independence Oracles On the Learning Curves of Revenue Maximization Asymptotically Robust Learning-Augmented Algorithms for Preemptive FIFO Buffer Management Flashback: A Reversible Bilateral Run-Peeling Decomposition of Strings Incremental Strongly Connected Components with Predictions Characterizing Admissible Objective Functions for Hierarchical Clustering Well-Conditioned Oblivious Perturbations in Linear Space Mathematical Foundations for Peer-to-Peer Lattice Computation Graph Neural Network-Informed Predictive Flows for Faster Ford-Fulkerson and PAC-Learnability A weighted angle distance on strings Towards Universal Convergence of Backward Error in Linear System Solvers Constant-Factor Approximations for Doubly Constrained Fair k-Center, k-Median and k-Means Tight Bounds for Learning Polyhedra with a Margin Efficiency of Proportional Mechanisms in Online Auto-Bidding Advertising Skyline-First Traversal as a Control Mechanism for Multi-Criteria Graph Search Constant-Factor Approximation for the Uniform Decision Tree Limited Perfect Monotonical Surrogates constructed using low-cost recursive linkage discovery with guaranteed output Query Lower Bounds for Diffusion Sampling Early Pruning for Public Transport Routing Adapting Dijkstra for Buffers and Unlimited Transfers Exploiting Low-Rank Structure in Max-K-Cut Problems Partial Optimality in the Preordering Problem High-accuracy log-concave sampling with stochastic queries Learning to Approximate Uniform Facility Location via Graph Neural Networks Linear Regression with Unknown Truncation Beyond Gaussian Features Adaptive Power Iteration Method for Differentially Private PCA Finite and Corruption-Robust Regret Bounds in Online Inverse Linear Optimization under M-Convex Action Sets Learning Mixture Models via Efficient High-dimensional Sparse Fourier Transforms Variance Computation for Weighted Model Counting with Knowledge Compilation Approach Deterministic Coreset for Lp Subspace Online Algorithms for Repeated Optimal Stopping: Balancing Baseline Guarantees and Regret Learned Static Function Data Structures Optimal hypersurface decision trees A Perfectly Truthful Calibration Measure The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm Best Agent Identification for General Game Playing A Faster Generalized Two-Stage Approximate Top-K Fast and Simple Densest Subgraph with Predictions Smoothed Analysis of Learning from Positive Samples Ineffectiveness for Search and Undecidability of PCSP Meta-Problems Sample-Efficient Optimization over Generative Priors via Coarse Learnability Efficient distributional regression trees learning algorithms for calibrated non-parametric probabilistic forecasts Testing Noise Assumptions of Learning Algorithms Testing Support Size More Efficiently Than Learning Histograms Sharper Bounds for Chebyshev Moment Matching, with Applications Expander Hierarchies for Normalized Cuts on Graphs Multilayer Correlation Clustering Efficient Parameter Estimation of Truncated Boolean Product Distributions Faster Hamiltonian Monte Carlo by Learning Leapfrog Scale: a self-calibrated randomized solution
Relative Error Fair Clustering in the Weak-Strong Oracle Model
Vladimir Braverman, Prathamesh Dharangutte, Shaofeng H. -C. Jian · 2025-06-14 · via cs.DS updates on arXiv.org

We study fair clustering problems in a setting where distance information is obtained from two sources: a strong oracle providing exact distances, but at a high cost, and a weak oracle providing potentially inaccurate distance estimates at a low cost. The goal is to produce a near-optimal fair clustering on $n$ input points with a minimum number of strong oracle queries. This models the increasingly common trade-off between accurate but expensive similarity measures (e.g., large-scale embeddings) and cheaper but inaccurate alternatives. The study of fair clustering in the model is motivated by the important quest of achieving fairness with the presence of inaccurate information. We achieve the first $(1+\varepsilon)$-coresets for fair $k$-median clustering using $\text{poly}\left(\frac{k}{\varepsilon}\cdot\log n\right)$ queries to the strong oracle. Furthermore, our results imply coresets for the standard setting (without fairness constraints), and we could in fact obtain $(1+\varepsilon)$-coresets for $(k,z)$-clustering for general $z=O(1)$ with a similar number of strong oracle queries. In contrast, previous results achieved a constant-factor $(>10)$ approximation for the standard $k$-clustering problems, and no previous work considered the fair $k$-median clustering problem.