Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,881 to 5,940 of 18,821

  1. Bernini: Latent Semantic Planning for Video Diffusion

    Bernini Team, Chenchen Liu, Junyi Chen +9

    cs.CVcs.AIcs.MMarXiv:2605.22344v12026
  2. Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

    Keming Wu, Zuhao Yang, Kaichen Zhang +24

    cs.CVarXiv:2604.28185v22026
  3. Linking spatial biology and clinical histology via Haiku

    Yan Cui, Jacob S. Leiby, Wenhui Lei +6

    cs.LGcs.CVq-bio.QMarXiv:2605.00925v12026
  4. Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models

    Fabian Morelli, Arnas Uselis, Ankit Sonthalia +1

    cs.CVarXiv:2605.15961v12026
  5. SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

    Kun Ouyang, Yuanxin Liu, Haoning Wu +5

    cs.CVarXiv:2504.01805v22025
  6. C-GenReg: Training-Free 3D Point Cloud Registration by Multi-View-Consistent Geometry-to-Image Generation with Probabilistic Modalities Fusion

    Yuval Haitman, Amit Efraim, Joseph M. Francos

    cs.CVarXiv:2604.16680v12026
  7. TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

    Bingyi Cao, Koert Chen, Kevis-Kokitsi Maninis +16

    cs.CVarXiv:2604.12012v12026
  8. Mario: Multimodal Graph Reasoning with Large Language Models

    Yuanfu Sun, Kang Li, Pengkang Guo +2

    cs.CVarXiv:2603.05181v22026
  9. UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?

    Zimo Wen, Boxiu Li, Wanbo Zhang +11

    cs.CVcs.AIarXiv:2603.03241v12026
  10. AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization

    Ashutosh Chaubey, Jiacheng Pang, Maksim Siniukov +1

    cs.LGcs.CVcs.HCarXiv:2602.07054v12026
  11. Generative Visual Code Mobile World Models

    Woosung Koh, Sungjun Han, Segyu Lee +2

    cs.LGcs.AIcs.CVarXiv:2602.01576v22026
  12. PISCO: Precise Video Instance Insertion with Sparse Control

    Xiangbo Gao, Renjie Li, Xinghao Chen +4

    cs.CVcs.AIarXiv:2602.08277v22026
  13. GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning

    Guoqing Ma, Siheng Wang, Zeyu Zhang +2

    cs.ROcs.CVarXiv:2602.04315v12026
  14. Finally Outshining the Random Baseline: A Simple and Effective Solution for Active Learning in 3D Biomedical Imaging

    Carsten T. Lüth, Jeremias Traub, Kim-Celine Kahl +6

    cs.CVarXiv:2601.13677v12026
  15. FAST-LIVO2: Fast, Direct LiDAR-Inertial-Visual Odometry

    Chunran Zheng, Wei Xu, Zuhao Zou +11

    cs.ROcs.CVarXiv:2408.14035v22024
  16. LION: Latent Point Diffusion Models for 3D Shape Generation

    Xiaohui Zeng, Arash Vahdat, Francis Williams +4

    cs.CVcs.LGstat.MLarXiv:2210.06978v12022
  17. AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

    Jun Zhan, Junqi Dai, Jiasheng Ye +13

    cs.CLcs.AIcs.CVarXiv:2402.12226v52024
  18. MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

    Xin Liu, Yichen Zhu, Jindong Gu +3

    cs.CVarXiv:2311.17600v52023
  19. Comparing deep neural networks against humans: object recognition when the signal gets weaker

    Robert Geirhos, David H. J. Janssen, Heiko H. Schütt +3

    cs.CVq-bio.NCstat.MLarXiv:1706.06969v22017
  20. Sign Language Recognition, Generation, and Translation: An Interdisciplinary Perspective

    Danielle Bragg, Oscar Koller, Mary Bellard +9

    cs.CVcs.CLcs.CYarXiv:1908.08597v12019
  21. Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

    Yaoting Wang, Shengqiong Wu, Yuecheng Zhang +4

    cs.CVarXiv:2503.12605v22025
  22. Automated Latent Fingerprint Recognition

    Kai Cao, Anil K. Jain

    cs.CVarXiv:1704.01925v12017
  23. Discovering Hidden Factors of Variation in Deep Networks

    Brian Cheung, Jesse A. Livezey, Arjun K. Bansal +1

    cs.LGcs.CVcs.NEarXiv:1412.6583v42014
  24. MAI-UI Technical Report: Real-World Centric Foundation GUI Agents

    Hanzhang Zhou, Xu Zhang, Panrong Tong +8

    cs.CVarXiv:2512.22047v12025
  25. Intriguing properties of synthetic images: from generative adversarial networks to diffusion models

    Riccardo Corvi, Davide Cozzolino, Giovanni Poggi +2

    cs.CVarXiv:2304.06408v22023
  26. Chained Predictions Using Convolutional Neural Networks

    Georgia Gkioxari, Alexander Toshev, Navdeep Jaitly

    cs.CVarXiv:1605.02346v22016
  27. Phantom: Subject-consistent video generation via cross-modal alignment

    Lijie Liu, Tianxiang Ma, Bingchuan Li +6

    cs.CVcs.AIarXiv:2502.11079v22025
  28. Magma: A Foundation Model for Multimodal AI Agents

    Jianwei Yang, Reuben Tan, Qianhui Wu +10

    cs.CVcs.AIcs.HCarXiv:2502.13130v12025
  29. Stabilizing Camera-Controlled Novel View Synthesis at Inference Time

    Prajwal Singh, Arjun Badola, Seema Kumari +2

    cs.CVarXiv:2609.03639v12026
  30. RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning

    Howard Qian, Yiting Chen, Yunfei Xie +6

    cs.CVcs.ROarXiv:2609.03199v12026
  31. Adversarial Generation of Continuous Images

    Ivan Skorokhodov, Savva Ignatyev, Mohamed Elhoseiny

    cs.CVcs.AIcs.LGarXiv:2011.12026v22020
  32. MMBERT: Multimodal BERT Pretraining for Improved Medical VQA

    Yash Khare, Viraj Bagal, Minesh Mathew +3

    cs.CVcs.CLcs.LGarXiv:2104.01394v12021
  33. Weighted Low-rank Tensor Recovery for Hyperspectral Image Restoration

    Yi Chang, Luxin Yan, Houzhang Fang +2

    cs.CVarXiv:1709.00192v12017
  34. Learning Plannable Representations with Causal InfoGAN

    Thanard Kurutach, Aviv Tamar, Ge Yang +2

    cs.LGcs.AIcs.CVarXiv:1807.09341v12018
  35. Revisiting Anchor Mechanisms for Temporal Action Localization

    Le Yang, Houwen Peng, Dingwen Zhang +2

    cs.CVarXiv:2008.09837v12020
  36. OCTA-500: A Retinal Dataset for Optical Coherence Tomography Angiography Study

    Mingchao Li, Kun Huang, Qiuzhuo Xu +7

    eess.IVcs.CVarXiv:2012.07261v32020
  37. From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers

    Jiacheng Liu, Chang Zou, Yuanhuiyi Lyu +2

    cs.CVcs.AIarXiv:2503.06923v22025
  38. Country-wide high-resolution vegetation height mapping with Sentinel-2

    Nico Lang, Konrad Schindler, Jan Dirk Wegner

    eess.IVcs.CVcs.LGarXiv:1904.13270v22019
  39. LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

    Shaolei Zhang, Qingkai Fang, Zhe Yang +1

    cs.CVcs.AIcs.CLarXiv:2501.03895v22025
  40. Deep Feature Space Trojan Attack of Neural Networks by Controlled Detoxification

    Siyuan Cheng, Yingqi Liu, Shiqing Ma +1

    cs.LGcs.CVarXiv:2012.11212v22020
  41. WorldScore: A Unified Evaluation Benchmark for World Generation

    Haoyi Duan, Hong-Xing Yu, Sirui Chen +2

    cs.GRcs.AIcs.CVarXiv:2504.00983v22025
  42. PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding

    Wei Chow, Jiageng Mao, Boyi Li +3

    cs.CVcs.AIcs.CLarXiv:2501.16411v22025
  43. Improved Anomaly Detection in Crowded Scenes via Cell-based Analysis of Foreground Speed, Size and Texture

    Vikas Reddy, Conrad Sanderson, Brian C. Lovell

    cs.CVarXiv:1304.0886v12013
  44. ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing

    Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer +7

    cs.LGcs.AIcs.CVarXiv:2608.26083v12026
  45. Domain-Size Pooling in Local Descriptors: DSP-SIFT

    Jingming Dong, Stefano Soatto

    cs.CVarXiv:1412.8556v32014
  46. Reconstructing Hand-Object Interactions in the Wild

    Zhe Cao, Ilija Radosavovic, Angjoo Kanazawa +1

    cs.CVarXiv:2012.09856v22020
  47. PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology

    Sheethal Bhat, Mahfuzur Rahman Chowdhury, Paula Andrea Perez-Toro +4

    cs.CVcs.AIarXiv:2608.25970v12026
  48. WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

    Jack Hong, Shilin Yan, Jiayin Cai +3

    cs.CVcs.AIarXiv:2502.04326v32025
  49. Less Contouring, More Accuracy: Lesion-Guided ROI Deep Learning for Ovarian Ultrasound Classification

    Mehran Ahmad, Ali Abbasian Ardakani, Afshin Mohammadi +3

    cs.CVarXiv:2608.25965v12026
  50. Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

    Yijun Yang, Shenghe Zheng, Wenbo Li +8

    cs.CVarXiv:2609.03729v12026
  51. RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition

    Xiaoyu Yue, Zhanghui Kuang, Chenhao Lin +2

    cs.CVarXiv:2007.07542v22020
  52. Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

    Haoyu Wang, Songchun Zhang, Haoran Li +3

    cs.CVcs.GRarXiv:2609.03557v12026
  53. GRIT: Teaching MLLMs to Think with Images

    Yue Fan, Xuehai He, Diji Yang +6

    cs.CVcs.AIcs.CLarXiv:2505.15879v22025
  54. Multiple Expert Brainstorming for Domain Adaptive Person Re-identification

    Yunpeng Zhai, Qixiang Ye, Shijian Lu +3

    cs.CVarXiv:2007.01546v32020
  55. Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

    Ye Wang, Ziheng Wang, Boshen Xu +14

    cs.CVcs.AIcs.CLarXiv:2503.13377v32025
  56. Metrics reloaded: Recommendations for image analysis validation

    Lena Maier-Hein, Annika Reinke, Patrick Godau +71

    cs.CVarXiv:2206.01653v82022
    Summaries:한국어
  57. Swarm-SLAM : Sparse Decentralized Collaborative Simultaneous Localization and Mapping Framework for Multi-Robot Systems

    Pierre-Yves Lajoie, Giovanni Beltrame

    cs.ROcs.CVarXiv:2301.06230v32023
  58. Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models

    Yuxiang Lai, Jike Zhong, Ming Li +4

    cs.CVarXiv:2503.13939v52025
  59. Naive-Deep Face Recognition: Touching the Limit of LFW Benchmark or Not?

    Erjin Zhou, Zhimin Cao, Qi Yin

    cs.CVarXiv:1501.04690v12015
  60. Human Mesh Recovery from Monocular Images via a Skeleton-disentangled Representation

    Sun Yu, Ye Yun, Liu Wu +3

    cs.CVarXiv:1908.07172v22019