惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Cisco Talos Blog
Cisco Talos Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Google Online Security Blog
Google Online Security Blog
博客园 - Franky
Hugging Face - Blog
Hugging Face - Blog
Security Archives - TechRepublic
Security Archives - TechRepublic
博客园 - 司徒正美
N
News and Events Feed by Topic
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
WordPress大学
WordPress大学
博客园 - 三生石上(FineUI控件)
Help Net Security
Help Net Security
N
News and Events Feed by Topic
O
OpenAI News
L
LangChain Blog
F
Full Disclosure
A
About on SuperTechFans
The GitHub Blog
The GitHub Blog
GbyAI
GbyAI
Cloudbric
Cloudbric
W
WeLiveSecurity
Application and Cybersecurity Blog
Application and Cybersecurity Blog
罗磊的独立博客
Attack and Defense Labs
Attack and Defense Labs
PCI Perspectives
PCI Perspectives
TaoSecurity Blog
TaoSecurity Blog
AI
AI
有赞技术团队
有赞技术团队
酷 壳 – CoolShell
酷 壳 – CoolShell
C
CXSECURITY Database RSS Feed - CXSecurity.com
C
Cisco Blogs
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Apple Machine Learning Research
Apple Machine Learning Research
C
CERT Recently Published Vulnerability Notes
T
The Exploit Database - CXSecurity.com
T
Threatpost
P
Palo Alto Networks Blog
G
GRAHAM CLULEY
Last Week in AI
Last Week in AI
雷峰网
雷峰网
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
C
Cyber Attacks, Cyber Crime and Cyber Security
博客园 - 聂微东
P
Proofpoint News Feed
Latest news
Latest news
S
SegmentFault 最新的问题
J
Java Code Geeks
T
Threat Research - Cisco Blogs
H
Help Net Security
P
Privacy International News Feed

Apple Machine Learning Research

LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning Environment-free Synthetic Data Generation for API-Calling Agents Accelerating Text-to-Video Generation with Calibrated Sparse Attention RayRoPE: Projective Ray Positional Encoding for Multi-View Attention LVSum: A Benchmark for Timestamp-Aware Long Video Summarization Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling When Unlearning Is Free: Leveraging Low Influence Points to Reduce Computational Costs Show Me Examples: Inferring Visual Concepts from Image Sets Location-Invariant Properties of Functions Versus Properties of Distributions: United in Testing but Separated in Verification Interactive Proofs for General Distribution Properties Doubly Sub-linear Interactive Proofs of Proximity Personalizing Incremental Video Search with Hybrid Text and ID Embeddings Embarrassingly Simple Self-Distillation Improves Code Generation CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning Uncertainty Quantification for LLM Function-Calling One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants Multilingual Semantic Retrieval for Apple Music Search Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies Incentivizing Temporal-Awareness in Egocentric Video Understanding Models Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction DynaMiCS: Fine-Tuning LLMs with Performance Constraints Using Dynamic Mixtures LensVLM: Selective Context Expansion for Compressed Visual Representation of Text Weblica: Scalable and Reproducible Training Environments for Visual Web Agents FlowEval: Reference-Based Evaluation of Generated User Interfaces A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models Scaling Properties of Continuous Diffusion Spoken Language Models Path-Constrained Mixture-of-Experts Revisiting ASR Error Correction with Specialized Models TopoPrimer: The Missing Topological Context in Forecasting Models Multi-Agent Teams Hold Experts Back VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization Amortizing Maximum Inner Product Search with Learned Support Functions On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers Learning Structured Reasoning via Tractable Trajectory Control Learning Unmasking Policies for Diffusion Language Models Residual Context Diffusion Language Models Conformal Thinking: Risk Control for Reasoning on a Compute Budget Anti-Causal Domain Generalization: Leveraging Unlabeled Data Metric-Dependent Annotation Saturation for Learning from Label Distributions Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels Introducing the Third Generation of Apple’s Foundation Models IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026 VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning Apple Workshop on Privacy-Preserving Machine Learning & AI 2026 Velox: Learning Representations of 4D Geometry and Appearance RVPO: Risk-Sensitive Alignment via Variance Regularization Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures Text-Conditional JEPA for Learning Semantically Rich Visual Representations What Matters in Practical Learned Image Compression SpecMD: A Comprehensive Study on Speculative Expert Prefetching From Where Things Are to What They’re For: Benchmarking Spatial–Functional Intelligence for Multimodal LLMs STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows Bootstrapping Sign Language Annotations with Sign Language Models International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2026 Adaptive Thinking: Large Language Models Know When to Think in Latent Space DSO: Direct Steering Optimization for Bias Mitigation StereoFoley: Object-Aware Stereo Audio Generation from Video LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning Local Mechanisms of Compositional Generalization in Conditional Diffusion Learning Long-Term Motion Embeddings for Efficient Kinematics Generation ParaRNN: Large-Scale Nonlinear RNNs, Trainable in Parallel Apple Machine Learning Research at ICLR 2026 Can Large Language Models Understand Context? International Conference on Learning Representations (ICLR) 2026 Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts Efficient Privacy Loss Accounting for Subsampling and Random Allocation ACM Human-Computer Interaction Conference (CHI) 2026 A Theoretical Framework for Acoustic Neighbor Embeddings Governance-Aware Agent Telemetry for Closed-Loop Enforcement in Multi-Agent AI Systems SQUIRE: Interactive UI Authoring via Slot QUery Intermediate REpresentations Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment ProText: A Benchmark Dataset for Measuring (Mis)gendering in Long-Form Texts Beyond Real Data: Synthetic Data through the Lens of Regularization Entropy-Preserving Reinforcement Learning Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting
MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching
2026-07-07 · via Apple Machine Learning Research

AuthorsJiahui Huang*, Yasi Zhang†*, Tianyu Chen‡, Shu Wang, Jianwen Xie§, Oscar Leong†, Mingyuan Zhou‡, Nanzhu Wang, Ying Nian Wu†

Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday users. However, editing models trained primarily for single-turn edits often break down in multi-turn editing—the natural interactive setting where a user iteratively refines an image based on the model’s own previous outputs. This failure stems from the all-or-nothing requirement, where a single failed turn compromises the entire sequence, and error propagation, where exposure bias leads to compounding editing errors. To address these challenges, we introduce MT-EditFlow, a flow-matching reinforcement learning framework designed to optimize reward signals for sequential image editing. MT-EditFlow integrates a multi-turn perspective with a multi-reward formulation to provide a unified structure applicable to both GRPO and NFT-based reinforcement learning methods. We systematically analyze and optimize the reward signal by investigating effective scoring strategies for turn-level aggregation, VLM reasoning modes to trade off reward bias and variance, and advantage fusion levels to prevent reward hacking. Our findings reveal that broadcasting the aggregated advantage across the entire editing trajectory effectively bridges the gap between local planning and global multi-turn task success. Extensive experiments demonstrate that MT-EditFlow significantly improves performance across diverse base models. Notably, it boosts FLUX.1-Kontext-dev by 6.85 points in turn-3 overall performance, surpassing state-of-the-art open-source models such as Qwen-Image-Edit. By maintaining high marginal success rates and reducing exposure bias, MT-EditFlow provides a foundation for more reliable and natural human-AI collaboration in visual content creation.

  • † University of California, Los Angeles
  • ‡ University of Texas at Austin
  • § Lambda, Inc
  • * Equal contribution

Related readings and updates.

We present UniGen-1.5, a unified multimodal large language model (MLLM) for advanced image understanding, generation and editing. Building upon UniGen, we comprehensively enhance the model architecture and training pipeline to strengthen the image understanding and generation capabilities while unlocking strong image editing ability. Especially, we propose a unified Reinforcement Learning (RL) strategy that improves both image generation and…

Read more

Recent advances in multimodal models have demonstrated remarkable text-guided image editing capabilities, with systems like GPT-4o and Nano-Banana setting new benchmarks. However, the research community’s progress remains constrained by the absence of large-scale, high-quality, and openly accessible datasets built from real images. We introduce Pico-Banana-400K, a comprehensive 400K-image dataset for instruction-based image editing. Our dataset…

Read more