Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
5,881 to 5,940 of 18,821
Bernini: Latent Semantic Planning for Video Diffusion
Bernini Team, Chenchen Liu, Junyi Chen +9
cs.CVcs.AIcs.MMarXiv:2605.22344v12026Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
Keming Wu, Zuhao Yang, Kaichen Zhang +24
cs.CVarXiv:2604.28185v22026Linking spatial biology and clinical histology via Haiku
Yan Cui, Jacob S. Leiby, Wenhui Lei +6
cs.LGcs.CVq-bio.QMarXiv:2605.00925v12026Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
Fabian Morelli, Arnas Uselis, Ankit Sonthalia +1
cs.CVarXiv:2605.15961v12026SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
Kun Ouyang, Yuanxin Liu, Haoning Wu +5
cs.CVarXiv:2504.01805v22025C-GenReg: Training-Free 3D Point Cloud Registration by Multi-View-Consistent Geometry-to-Image Generation with Probabilistic Modalities Fusion
Yuval Haitman, Amit Efraim, Joseph M. Francos
cs.CVarXiv:2604.16680v12026TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
Bingyi Cao, Koert Chen, Kevis-Kokitsi Maninis +16
cs.CVarXiv:2604.12012v12026Mario: Multimodal Graph Reasoning with Large Language Models
Yuanfu Sun, Kang Li, Pengkang Guo +2
cs.CVarXiv:2603.05181v22026UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
Zimo Wen, Boxiu Li, Wanbo Zhang +11
cs.CVcs.AIarXiv:2603.03241v12026AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization
Ashutosh Chaubey, Jiacheng Pang, Maksim Siniukov +1
cs.LGcs.CVcs.HCarXiv:2602.07054v12026Generative Visual Code Mobile World Models
Woosung Koh, Sungjun Han, Segyu Lee +2
cs.LGcs.AIcs.CVarXiv:2602.01576v22026PISCO: Precise Video Instance Insertion with Sparse Control
Xiangbo Gao, Renjie Li, Xinghao Chen +4
cs.CVcs.AIarXiv:2602.08277v22026GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
Guoqing Ma, Siheng Wang, Zeyu Zhang +2
cs.ROcs.CVarXiv:2602.04315v12026Finally Outshining the Random Baseline: A Simple and Effective Solution for Active Learning in 3D Biomedical Imaging
Carsten T. Lüth, Jeremias Traub, Kim-Celine Kahl +6
cs.CVarXiv:2601.13677v12026FAST-LIVO2: Fast, Direct LiDAR-Inertial-Visual Odometry
Chunran Zheng, Wei Xu, Zuhao Zou +11
cs.ROcs.CVarXiv:2408.14035v22024LION: Latent Point Diffusion Models for 3D Shape Generation
Xiaohui Zeng, Arash Vahdat, Francis Williams +4
cs.CVcs.LGstat.MLarXiv:2210.06978v12022AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Jun Zhan, Junqi Dai, Jiasheng Ye +13
cs.CLcs.AIcs.CVarXiv:2402.12226v52024MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
Xin Liu, Yichen Zhu, Jindong Gu +3
cs.CVarXiv:2311.17600v52023Comparing deep neural networks against humans: object recognition when the signal gets weaker
Robert Geirhos, David H. J. Janssen, Heiko H. Schütt +3
cs.CVq-bio.NCstat.MLarXiv:1706.06969v22017Sign Language Recognition, Generation, and Translation: An Interdisciplinary Perspective
Danielle Bragg, Oscar Koller, Mary Bellard +9
cs.CVcs.CLcs.CYarXiv:1908.08597v12019Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
Yaoting Wang, Shengqiong Wu, Yuecheng Zhang +4
cs.CVarXiv:2503.12605v22025Automated Latent Fingerprint Recognition
Kai Cao, Anil K. Jain
cs.CVarXiv:1704.01925v12017Discovering Hidden Factors of Variation in Deep Networks
Brian Cheung, Jesse A. Livezey, Arjun K. Bansal +1
cs.LGcs.CVcs.NEarXiv:1412.6583v42014MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
Hanzhang Zhou, Xu Zhang, Panrong Tong +8
cs.CVarXiv:2512.22047v12025Intriguing properties of synthetic images: from generative adversarial networks to diffusion models
Riccardo Corvi, Davide Cozzolino, Giovanni Poggi +2
cs.CVarXiv:2304.06408v22023Chained Predictions Using Convolutional Neural Networks
Georgia Gkioxari, Alexander Toshev, Navdeep Jaitly
cs.CVarXiv:1605.02346v22016Phantom: Subject-consistent video generation via cross-modal alignment
Lijie Liu, Tianxiang Ma, Bingchuan Li +6
cs.CVcs.AIarXiv:2502.11079v22025Magma: A Foundation Model for Multimodal AI Agents
Jianwei Yang, Reuben Tan, Qianhui Wu +10
cs.CVcs.AIcs.HCarXiv:2502.13130v12025Stabilizing Camera-Controlled Novel View Synthesis at Inference Time
Prajwal Singh, Arjun Badola, Seema Kumari +2
cs.CVarXiv:2609.03639v12026RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning
Howard Qian, Yiting Chen, Yunfei Xie +6
cs.CVcs.ROarXiv:2609.03199v12026Adversarial Generation of Continuous Images
Ivan Skorokhodov, Savva Ignatyev, Mohamed Elhoseiny
cs.CVcs.AIcs.LGarXiv:2011.12026v22020MMBERT: Multimodal BERT Pretraining for Improved Medical VQA
Yash Khare, Viraj Bagal, Minesh Mathew +3
cs.CVcs.CLcs.LGarXiv:2104.01394v12021Weighted Low-rank Tensor Recovery for Hyperspectral Image Restoration
Yi Chang, Luxin Yan, Houzhang Fang +2
cs.CVarXiv:1709.00192v12017Learning Plannable Representations with Causal InfoGAN
Thanard Kurutach, Aviv Tamar, Ge Yang +2
cs.LGcs.AIcs.CVarXiv:1807.09341v12018Revisiting Anchor Mechanisms for Temporal Action Localization
Le Yang, Houwen Peng, Dingwen Zhang +2
cs.CVarXiv:2008.09837v12020OCTA-500: A Retinal Dataset for Optical Coherence Tomography Angiography Study
Mingchao Li, Kun Huang, Qiuzhuo Xu +7
eess.IVcs.CVarXiv:2012.07261v32020From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers
Jiacheng Liu, Chang Zou, Yuanhuiyi Lyu +2
cs.CVcs.AIarXiv:2503.06923v22025Country-wide high-resolution vegetation height mapping with Sentinel-2
Nico Lang, Konrad Schindler, Jan Dirk Wegner
eess.IVcs.CVcs.LGarXiv:1904.13270v22019LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Shaolei Zhang, Qingkai Fang, Zhe Yang +1
cs.CVcs.AIcs.CLarXiv:2501.03895v22025Deep Feature Space Trojan Attack of Neural Networks by Controlled Detoxification
Siyuan Cheng, Yingqi Liu, Shiqing Ma +1
cs.LGcs.CVarXiv:2012.11212v22020WorldScore: A Unified Evaluation Benchmark for World Generation
Haoyi Duan, Hong-Xing Yu, Sirui Chen +2
cs.GRcs.AIcs.CVarXiv:2504.00983v22025PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
Wei Chow, Jiageng Mao, Boyi Li +3
cs.CVcs.AIcs.CLarXiv:2501.16411v22025Improved Anomaly Detection in Crowded Scenes via Cell-based Analysis of Foreground Speed, Size and Texture
Vikas Reddy, Conrad Sanderson, Brian C. Lovell
cs.CVarXiv:1304.0886v12013ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing
Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer +7
cs.LGcs.AIcs.CVarXiv:2608.26083v12026Domain-Size Pooling in Local Descriptors: DSP-SIFT
Jingming Dong, Stefano Soatto
cs.CVarXiv:1412.8556v32014Reconstructing Hand-Object Interactions in the Wild
Zhe Cao, Ilija Radosavovic, Angjoo Kanazawa +1
cs.CVarXiv:2012.09856v22020PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology
Sheethal Bhat, Mahfuzur Rahman Chowdhury, Paula Andrea Perez-Toro +4
cs.CVcs.AIarXiv:2608.25970v12026WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
Jack Hong, Shilin Yan, Jiayin Cai +3
cs.CVcs.AIarXiv:2502.04326v32025Less Contouring, More Accuracy: Lesion-Guided ROI Deep Learning for Ovarian Ultrasound Classification
Mehran Ahmad, Ali Abbasian Ardakani, Afshin Mohammadi +3
cs.CVarXiv:2608.25965v12026Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
Yijun Yang, Shenghe Zheng, Wenbo Li +8
cs.CVarXiv:2609.03729v12026RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition
Xiaoyu Yue, Zhanghui Kuang, Chenhao Lin +2
cs.CVarXiv:2007.07542v22020Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation
Haoyu Wang, Songchun Zhang, Haoran Li +3
cs.CVcs.GRarXiv:2609.03557v12026GRIT: Teaching MLLMs to Think with Images
Yue Fan, Xuehai He, Diji Yang +6
cs.CVcs.AIcs.CLarXiv:2505.15879v22025Multiple Expert Brainstorming for Domain Adaptive Person Re-identification
Yunpeng Zhai, Qixiang Ye, Shijian Lu +3
cs.CVarXiv:2007.01546v32020Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Ye Wang, Ziheng Wang, Boshen Xu +14
cs.CVcs.AIcs.CLarXiv:2503.13377v32025Metrics reloaded: Recommendations for image analysis validation
Lena Maier-Hein, Annika Reinke, Patrick Godau +71
cs.CVarXiv:2206.01653v82022Summaries:한국어Swarm-SLAM : Sparse Decentralized Collaborative Simultaneous Localization and Mapping Framework for Multi-Robot Systems
Pierre-Yves Lajoie, Giovanni Beltrame
cs.ROcs.CVarXiv:2301.06230v32023Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models
Yuxiang Lai, Jike Zhong, Ming Li +4
cs.CVarXiv:2503.13939v52025Naive-Deep Face Recognition: Touching the Limit of LFW Benchmark or Not?
Erjin Zhou, Zhimin Cao, Qi Yin
cs.CVarXiv:1501.04690v12015Human Mesh Recovery from Monocular Images via a Skeleton-disentangled Representation
Sun Yu, Ye Yun, Liu Wu +3
cs.CVarXiv:1908.07172v22019