惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Schneier on Security
Schneier on Security
N
Netflix TechBlog - Medium
IT之家
IT之家
MongoDB | Blog
MongoDB | Blog
博客园_首页
S
SegmentFault 最新的问题
H
Help Net Security
P
Proofpoint News Feed
云风的 BLOG
云风的 BLOG
T
The Blog of Author Tim Ferriss
量子位
GbyAI
GbyAI
M
MIT News - Artificial intelligence
Recorded Future
Recorded Future
P
Privacy & Cybersecurity Law Blog
B
Blog
月光博客
月光博客
博客园 - 聂微东
Vercel News
Vercel News
罗磊的独立博客
腾讯CDC
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
A
Arctic Wolf
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Stack Overflow Blog
Stack Overflow Blog
T
Threat Research - Cisco Blogs
Blog — PlanetScale
Blog — PlanetScale
L
Lohrmann on Cybersecurity
I
Intezer
小众软件
小众软件
T
The Exploit Database - CXSecurity.com
Jina AI
Jina AI
C
Check Point Blog
AWS News Blog
AWS News Blog
C
Cisco Blogs
Martin Fowler
Martin Fowler
The Last Watchdog
The Last Watchdog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
宝玉的分享
宝玉的分享
S
Security Affairs
大猫的无限游戏
大猫的无限游戏
N
News and Events Feed by Topic
雷峰网
雷峰网
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
H
Hacker News: Front Page
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
F
Full Disclosure
P
Proofpoint News Feed
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Microsoft Security Blog
Microsoft Security Blog

Apple Machine Learning Research

LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning Environment-free Synthetic Data Generation for API-Calling Agents Accelerating Text-to-Video Generation with Calibrated Sparse Attention RayRoPE: Projective Ray Positional Encoding for Multi-View Attention LVSum: A Benchmark for Timestamp-Aware Long Video Summarization Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling When Unlearning Is Free: Leveraging Low Influence Points to Reduce Computational Costs Location-Invariant Properties of Functions Versus Properties of Distributions: United in Testing but Separated in Verification Interactive Proofs for General Distribution Properties Doubly Sub-linear Interactive Proofs of Proximity Personalizing Incremental Video Search with Hybrid Text and ID Embeddings Embarrassingly Simple Self-Distillation Improves Code Generation CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning Uncertainty Quantification for LLM Function-Calling One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants Multilingual Semantic Retrieval for Apple Music Search Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies Incentivizing Temporal-Awareness in Egocentric Video Understanding Models Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction DynaMiCS: Fine-Tuning LLMs with Performance Constraints Using Dynamic Mixtures LensVLM: Selective Context Expansion for Compressed Visual Representation of Text MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching Weblica: Scalable and Reproducible Training Environments for Visual Web Agents FlowEval: Reference-Based Evaluation of Generated User Interfaces A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models Scaling Properties of Continuous Diffusion Spoken Language Models Path-Constrained Mixture-of-Experts Revisiting ASR Error Correction with Specialized Models TopoPrimer: The Missing Topological Context in Forecasting Models Multi-Agent Teams Hold Experts Back VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization Amortizing Maximum Inner Product Search with Learned Support Functions On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers Learning Structured Reasoning via Tractable Trajectory Control Learning Unmasking Policies for Diffusion Language Models Residual Context Diffusion Language Models Conformal Thinking: Risk Control for Reasoning on a Compute Budget Anti-Causal Domain Generalization: Leveraging Unlabeled Data Metric-Dependent Annotation Saturation for Learning from Label Distributions Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels Introducing the Third Generation of Apple’s Foundation Models IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026 VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning Apple Workshop on Privacy-Preserving Machine Learning & AI 2026 Velox: Learning Representations of 4D Geometry and Appearance RVPO: Risk-Sensitive Alignment via Variance Regularization Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures Text-Conditional JEPA for Learning Semantically Rich Visual Representations What Matters in Practical Learned Image Compression SpecMD: A Comprehensive Study on Speculative Expert Prefetching From Where Things Are to What They’re For: Benchmarking Spatial–Functional Intelligence for Multimodal LLMs STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows Bootstrapping Sign Language Annotations with Sign Language Models International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2026 Adaptive Thinking: Large Language Models Know When to Think in Latent Space DSO: Direct Steering Optimization for Bias Mitigation StereoFoley: Object-Aware Stereo Audio Generation from Video LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning Local Mechanisms of Compositional Generalization in Conditional Diffusion Learning Long-Term Motion Embeddings for Efficient Kinematics Generation ParaRNN: Large-Scale Nonlinear RNNs, Trainable in Parallel Apple Machine Learning Research at ICLR 2026 Can Large Language Models Understand Context? International Conference on Learning Representations (ICLR) 2026 Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts Efficient Privacy Loss Accounting for Subsampling and Random Allocation ACM Human-Computer Interaction Conference (CHI) 2026 A Theoretical Framework for Acoustic Neighbor Embeddings Governance-Aware Agent Telemetry for Closed-Loop Enforcement in Multi-Agent AI Systems SQUIRE: Interactive UI Authoring via Slot QUery Intermediate REpresentations Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment ProText: A Benchmark Dataset for Measuring (Mis)gendering in Long-Form Texts Beyond Real Data: Synthetic Data through the Lens of Regularization Entropy-Preserving Reinforcement Learning Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting
Show Me Examples: Inferring Visual Concepts from Image Sets
2026-07-17 · via Apple Machine Learning Research

AuthorsNick Stracke†, Kolja Bauer†, Josh Susskind, Miguel Angel Bautista Martin, Björn Ommer†

Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current models fail to infer shared concepts from sets of example images and apply them to new inputs. We introduce Visual Concept Inference from Sets (VICIS), a task that evaluates this capability. Given a small context set of images sharing a concept and a query image, the model must generate new images that preserve the context-defined concept while remaining consistent with the query. We show that state-of-the-art VLMs perform poorly on this task, often ignoring the visual context or defaulting to biased generations. To address this gap, we propose a training framework and architecture that learn to infer visual concepts from image sets and extract concept-specific embeddings from queries. Experiments on synthetic data and large-scale ImageNet/WordNet data show that our model generates more accurate and diverse outputs and generalizes to unseen concepts and modalities such as sketches.

  • † LMU

Related readings and updates.

In this work we study the presence of expert units in pre-trained Transformer Models (TM), and how they impact a model’s performance. We define expert units to be neurons that are able to classify a concept with a given average precision, where a concept is represented by a binary set of sentences containing the concept (or not). Leveraging the OneSec dataset (Scarlini et al., 2019), we compile a dataset of 1641 concepts that allows diverse…

Read more

Most successful examples of neural nets today are trained with supervision. However, to achieve high accuracy, the training sets need to be large, diverse, and accurately annotated, which is costly. An alternative to labelling huge amounts of data is to use synthetic images from a simulator. This is cheap as there is no labeling cost, but the synthetic images may not be realistic enough, resulting in poor generalization on real test images. To help close this performance gap, we’ve developed a method for refining synthetic images to make them look more realistic. We show that training models on these refined images leads to significant improvements in accuracy on various machine learning tasks.

Read more