Planning with Unified Multimodal Models - 惯性聚合

推荐订阅源

Darknet – Hacking Tools, Hacker News & Cyber Security

KPMG report finds enterprise disconnect between AI and its ROI | CIO

Schneier on Security

Secure Thoughts

Security Archives - TechRepublic

Exploit-DB.com RSS Feed

The Hacker News

Know Your Adversary

Threat Intelligence Blog | Flashpoint

Kaspersky official blog

Forbes - Security

TaoSecurity Blog

Simon Willison's Weblog

Cyber Attacks, Cyber Crime and Cyber Security

Vulnerabilities – Threatpost

DataBreaches.Net

The Last Watchdog

CTFtime.org: upcoming CTF events

Palo Alto Networks Blog

钛媒体：引领未来商业与生活新知

CERT Recently Published Vulnerability Notes

The Register - Security

Stack Overflow Blog

Microsoft Azure Blog

Hackread – Cybersecurity News, Data Breaches, AI and More

Hacker News: Front Page

Recorded Future

Fortinet All Blogs

大猫的无限游戏

cs.AI updates on arXiv.org

About on SuperTechFans

Privacy International News Feed

Tailwind CSS Blog

Privacy & Cybersecurity Law Blog

cs.CV updates on arXiv.org

A Lightweight Multi-Metric No-Reference Image Quality Assessment Framework for UAV Imaging PatchPoison: Poisoning Multi-View Datasets to Degrade 3D Reconstruction 3DRealHead: Few-Shot Detailed Head Avatar GeoLink: A 3D-Aware Framework Towards Better Generalization in Cross-View Geo-Localization Towards Patient-Specific Deformable Registration in Laparoscopic Surgery Neural 3D Reconstruction of Planetary Surfaces from Descent-Phase Wide-Angle Imagery A High-Resolution Landscape Dataset for Concept-Based XAI With Application to Species Distribution Models DroneScan-YOLO: Redundancy-Aware Lightweight Detection for Tiny Objects in UAV Imagery See&Say: Vision Language Guided Safe Zone Detection for Autonomous Package Delivery Drones PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines Bias at the End of the Score Deep Spatially-Regularized and Superpixel-Based Diffusion Learning for Unsupervised Hyperspectral Image Clustering The Spectrascapes Dataset: Street-view imagery beyond the visible captured using a mobile platform Why MLLMs Struggle to Determine Object Orientations Towards Successful Implementation of Automated Raveling Detection: Effects of Training Data Size, Illumination Difference, and Spatial Shift Right Regions, Wrong Labels: Semantic Label Flips in Segmentation under Correlation Shift Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision Explainable Fall Detection for Elderly Monitoring via Temporally Stable SHAP in Skeleton-Based Human Activity Recognition Indexing Multimodal Language Models for Large-scale Image Retrieval Rethinking Uncertainty in Segmentation: From Estimation to Decision 4th Workshop on Maritime Computer Vision (MaCVi): Challenge Overview SemiFA: An Agentic Multi-Modal Framework for Autonomous Semiconductor Failure Analysis Report Generation Multitasking Embedding for Embryo Blastocyst Grading Prediction (MEmEBG) Graph Propagated Projection Unlearning: A Unified Framework for Vision and Audio Discriminative Models Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference Anthropogenic Regional Adaptation in Multimodal Vision-Language Model ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation Lightweight Low-Light Image Enhancement via Distribution-Normalizing Preprocessing and Depthwise U-Net You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass A Semi-Automated Framework for 3D Reconstruction of Medieval Manuscript Miniatures ViSAGE @ NTIRE 2026 Challenge on Video Saliency Prediction InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation EfficientSign: An Attention-Enhanced Lightweight Architecture for Indian Sign Language Recognition Unified Multimodal Uncertain Inference State Space Models are Effective Sign Language Learners: Exploiting Phonological Compositionality for Vocabulary-Scale Recognition Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models CatalogStitch: Dimension-Aware and Occlusion-Preserving Object Compositing for Catalog Image Generation DeFakeQ: Enabling Real-Time Deepfake Detection on Edge Devices via Adaptive Bidirectional Quantization BIAS: A Biologically Inspired Algorithm for Video Saliency Detection Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation Scenarios Dynamic Class-Aware Active Learning for Unbiased Satellite Image Segmentation Domain-generalizable Face Anti-Spoofing with Patch-based Multi-tasking and Artifact Pattern Conversion Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation Detecting Diffusion-generated Images via Dynamic Assembly Forests CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation Long-SCOPE: Fully Sparse Long-Range Cooperative 3D Perception Adding Another Dimension to Image-based Animal Detection Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search Through Their Eyes: Fixation-aligned Tuning for Personalized User Emulation FDIF: Formula-Driven supervised Learning with Implicit Functions for 3D Medical Image Segmentation B-MoE: A Body-Part-Aware Mixture-of-Experts "All Parts Matter" Approach to Micro-Action Recognition ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing COREY: Entropy-Guided Runtime Chunk Scheduling for Selective Scan Kernels Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex Degradation-Consistent Paired Training for Robust AI-Generated Image Detection Genie 4D: Semantic-Prior-Guided 4D Dynamic Scene Reconstruction Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories Neural Distribution Prior for LiDAR Out-of-Distribution Detection Adaptive Dual Residual U-Net with Attention Gate and Multiscale Spatial Attention Mechanisms (ADRUwAMS) SenBen: Sensitive Scene Graphs for Explainable Content Moderation Unsupervised Local Plasticity in a Multi-Frequency VisNet Hierarchy 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding On Semiotic-Grounded Interpretive Evaluation of Generative Art Detection of Hate and Threat in Digital Forensics: A Case-Driven Multimodal Approach VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG StableTTA: Improving Vision Model Performance by Training-free Test-Time Adaptation Methods Grid2Matrix: Revealing Digital Agnosia in Vision-Language Models Belief-Aware VLM Model for Human-like Reasoning Zero-Shot Quantization via Weight-Space Arithmetic Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behavior VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation DecepGPT: Schema-Driven Deception Detection with Multicultural Datasets and Robust Multimodal Learning Interpretable Alzheimer's Diagnosis via Multimodal Fusion of Regional Brain Experts Geometry-Aware Cross Modal Alignment for Light Field-LiDAR Semantic Segmentation PnP-CM: Consistency Models as Plug-and-Play Priors for Inverse Problems KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation FedKLPR: KL-Guided Pruning-Aware Federated Learning for Person Re-Identification COXNet: Cross-Layer Fusion with Adaptive Alignment and Scale Integration for RGBT Tiny Object Detection AdvDINO: Domain-Adversarial Self-Supervised Representation Learning for Spatial Proteomics PRIX: Learning to Plan from Raw Pixels for End-to-End Autonomous Driving Progressive Multimodal Interaction Network for Reliable Quantification of Fish Feeding Intensity in Aquaculture VRAG: Learning World Models for Interactive Video Generation GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping Variational Visual Question Answering for Uncertainty-Aware Selective Prediction Auto-regressive transformation for image alignment LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection AccidentSim: Generating Vehicle Collision Videos with Physically Realistic Collision Trajectories from Real-World Accident Reports Integrating Semi-Supervised and Active Learning for Semantic Segmentation HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks OmniPrism: Learning Disentangled Visual Concept for Image Generation Linear Attention Based Deep Nonlocal Means Filtering for Multiplicative Noise Removal MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions

Planning with Unified Multimodal Models

Yihao Sun, Zhilong Zhang, Yang Yu, Pierre-Luc Bacon · 2025-09-27 · via cs.CV updates on arXiv.org

此内容由惯性聚合(RSS阅读器)自动聚合整理，仅供阅读参考。原文来自 — 版权归原作者所有。