Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

15,421 to 15,480 of 18,821

  1. AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

    Xilin Jiang, Qiaolin Wang, Junkai Wu +30

    cs.SDcs.CLcs.CVarXiv:2601.17645v12026
  2. SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

    Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum +2

    cs.AIcs.CVarXiv:2601.21666v22026
  3. Videos as Space-Time Region Graphs

    Xiaolong Wang, Abhinav Gupta

    cs.CVarXiv:1806.01810v22018
  4. DeepIM: Deep Iterative Matching for 6D Pose Estimation

    Yi Li, Gu Wang, Xiangyang Ji +2

    cs.CVcs.ROarXiv:1804.00175v42018
  5. End-to-End Learning of Visual Representations from Uncurated Instructional Videos

    Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira +3

    cs.CVarXiv:1912.06430v42019
  6. HyperAlign: Hypernetwork for Efficient Test-Time Alignment of Diffusion Models

    Xin Xie, Jiaxian Guo, Dong Gong

    cs.CVarXiv:2601.15968v22026
  7. Embodied Question Answering

    Abhishek Das, Samyak Datta, Georgia Gkioxari +3

    cs.CVcs.AIcs.CLarXiv:1711.11543v22017
  8. DeFM: Learning Foundation Representations from Depth for Robotics

    Manthan Patel, Jonas Frey, Mayank Mittal +5

    cs.ROcs.CVarXiv:2601.18923v12026
  9. Quality Inspection of Printed Circuit Board Pin Insertion via Semantic Segmentation and Board-Level Feature Extraction

    Nils Rabeneck, André Kiunke, Nicole Hoess +1

    cs.CVarXiv:2608.22937v12026
  10. Novel methods for multilinear data completion and de-noising based on tensor-SVD

    Zemin Zhang, Gregory Ely, Shuchin Aeron +2

    cs.CVarXiv:1407.1785v22014
  11. Towards Pixel-Level VLM Perception via Simple Points Prediction

    Tianhui Song, Haoyu Lu, Hao Yang +8

    cs.CVarXiv:2601.19228v12026
  12. Hyperspherical Autoencoder for High-Fidelity Image Reconstruction and Generation

    Hun Chang, Byunghee Cha, Jong Chul Ye

    cs.CVcs.AIcs.LGarXiv:2601.22904v22026
  13. Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning

    Yu Xu, Yuxin Zhang, Juan Cao +5

    cs.CVcs.AIarXiv:2602.01335v12026
  14. DRIT++: Diverse Image-to-Image Translation via Disentangled Representations

    Hsin-Ying Lee, Hung-Yu Tseng, Qi Mao +4

    cs.CVarXiv:1905.01270v22019
  15. SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-body Manipulation

    Mu Huang, Hui Wang, Kerui Ren +5

    cs.ROcs.AIcs.CVarXiv:2602.02402v22026
  16. Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers

    Bozhou Li, Yushuo Guan, Haolin Li +7

    cs.CVarXiv:2602.03510v12026
  17. Delving Deeper into Convolutional Networks for Learning Video Representations

    Nicolas Ballas, Li Yao, Chris Pal +1

    cs.CVcs.LGcs.NEarXiv:1511.06432v42015
  18. Capsule-Forensics: Using Capsule Networks to Detect Forged Images and Videos

    Huy H. Nguyen, Junichi Yamagishi, Isao Echizen

    cs.CVeess.IVarXiv:1810.11215v12018
  19. CGNet: A Light-weight Context Guided Network for Semantic Segmentation

    Tianyi Wu, Sheng Tang, Rui Zhang +1

    cs.CVarXiv:1811.08201v22018
  20. Scaling Up Your Kernels to 31x31: Revisiting Large Kernel Design in CNNs

    Xiaohan Ding, Xiangyu Zhang, Yizhuang Zhou +3

    cs.CVcs.AIcs.LGarXiv:2203.06717v42022
  21. TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving

    Kashyap Chitta, Aditya Prakash, Bernhard Jaeger +3

    cs.CVcs.AIcs.LGarXiv:2205.15997v12022
  22. Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks

    Bohan Zeng, Kaixin Zhu, Daili Hua +24

    cs.CVarXiv:2602.01630v12026
  23. Working hard to know your neighbor's margins: Local descriptor learning loss

    Anastasiya Mishchuk, Dmytro Mishkin, Filip Radenovic +1

    cs.CVarXiv:1705.10872v42017
  24. 3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation

    Zhixue Fang, Xu He, Songlin Tang +5

    cs.CVarXiv:2602.03796v22026
  25. Attribute2Image: Conditional Image Generation from Visual Attributes

    Xinchen Yan, Jimei Yang, Kihyuk Sohn +1

    cs.LGcs.AIcs.CVarXiv:1512.00570v22015
  26. How Merge-Tolerant Are Vision Transformers for Wheat Phenotyping?

    Simon Ravé, Pejman Rasti, David Rousseau

    cs.CVarXiv:2608.23142v12026
  27. TransTrack: Multiple Object Tracking with Transformer

    Peize Sun, Jinkun Cao, Yi Jiang +5

    cs.CVarXiv:2012.15460v22020
  28. Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields

    Dor Verbin, Peter Hedman, Ben Mildenhall +3

    cs.CVcs.GRarXiv:2112.03907v12021
  29. Adaptive 1D Video Diffusion Autoencoder

    Yao Teng, Minxuan Lin, Xian Liu +3

    cs.CVarXiv:2602.04220v12026
  30. When the Edit Changes the Patient: Measuring Identity Preservation in Counterfactual Retinal Images

    Andrea Posada, Wenke Karbole, Bach Ngoc Doan +9

    cs.CVarXiv:2608.23024v12026
  31. OmniRad: A Radiological Foundation Model for Multi-Task Medical Image Analysis

    Luca Zedda, Andrea Loddo, Cecilia Di Ruberto

    cs.CVcs.AIarXiv:2602.04547v12026
  32. No More Strided Convolutions or Pooling: A New CNN Building Block for Low-Resolution Images and Small Objects

    Raja Sunkara, Tie Luo

    cs.CVcs.LGarXiv:2208.03641v12022
  33. Cross Attention Network for Few-shot Classification

    Ruibing Hou, Hong Chang, Bingpeng Ma +2

    cs.CVarXiv:1910.07677v12019
  34. Human Action Recognition from Various Data Modalities: A Review

    Zehua Sun, Qiuhong Ke, Hossein Rahmani +3

    cs.CVarXiv:2012.11866v52020
  35. When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models

    Jiacheng Hou, Yining Sun, Ruochong Jin +4

    cs.CVcs.AIarXiv:2602.10179v22026
  36. Object Goal Navigation using Goal-Oriented Semantic Exploration

    Devendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta +1

    cs.CVcs.LGcs.ROarXiv:2007.00643v22020
  37. Neighbor-Aware View Synthesis for Restoring Missing Views in Light-Field Camera Arrays

    Sakshi Goel, Ayush Goyal, K S Venkatesh +1

    cs.CVarXiv:2608.23175v12026
  38. TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution

    Deyang Jiang, Jing Huang, Xuanle Zhao +6

    cs.CVarXiv:2602.09662v12026
  39. DeepFakes: a New Threat to Face Recognition? Assessment and Detection

    Pavel Korshunov, Sebastien Marcel

    cs.CVarXiv:1812.08685v12018
  40. Reinforced Attention Learning

    Bangzheng Li, Jianmo Ni, Chen Qu +5

    cs.CLcs.CVcs.LGarXiv:2602.04884v22026
  41. Auto-Encoding Scene Graphs for Image Captioning

    Xu Yang, Kaihua Tang, Hanwang Zhang +1

    cs.CVarXiv:1812.02378v32018
  42. Fast-SAM3D: 3Dfy Anything in Images but Faster

    Weilun Feng, Mingqiang Wu, Zhiliang Chen +10

    cs.CVarXiv:2602.05293v22026
  43. Prism: Spectral-Aware Block-Sparse Attention

    Xinghao Wang, Pengyu Wang, Xiaoran Liu +4

    cs.CLcs.AIcs.CVarXiv:2602.08426v22026
  44. MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE

    Ruijie Zhu, Jiahao Lu, Wenbo Hu +4

    cs.CVcs.AIcs.CGarXiv:2602.08961v22026
  45. ArcFlow: Unleashing 2-Step Text-to-Image Generation via High-Precision Non-Linear Flow Distillation

    Zihan Yang, Shuyuan Tu, Licheng Zhang +3

    cs.CVcs.AIarXiv:2602.09014v12026
  46. Joint Optimization Framework for Learning with Noisy Labels

    Daiki Tanaka, Daiki Ikami, Toshihiko Yamasaki +1

    cs.CVcs.LGstat.MLarXiv:1803.11364v12018
  47. Reasoning-Augmented Representations for Multimodal Retrieval

    Jianrui Zhang, Anirudh Sundara Rajan, Brandon Han +3

    cs.IRcs.AIcs.CVarXiv:2602.07125v12026
  48. VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval

    Issar Tzachor, Dvir Samuel, Rami Ben-Ari

    cs.CVcs.AIarXiv:2602.08099v12026
  49. Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning

    Xu Ma, Yitian Zhang, Qihua Dong +1

    cs.CVarXiv:2602.09439v12026
  50. Motion-Based Tokenization for Cross-Dataset Egocentric Gaze Modeling

    Virmarie Maquiling, Zhuojiang Cai, Enkelejda Kasneci

    cs.CVarXiv:2608.22926v12026
  51. Video Frame Synthesis using Deep Voxel Flow

    Ziwei Liu, Raymond A. Yeh, Xiaoou Tang +2

    cs.CVcs.GRcs.LGarXiv:1702.02463v22017
  52. Gated Siamese Convolutional Neural Network Architecture for Human Re-Identification

    Rahul Rama Varior, Mrinal Haloi, Gang Wang

    cs.CVarXiv:1607.08378v22016
  53. Multi-Task Learning with Deep Neural Networks: A Survey

    Michael Crawshaw

    cs.LGcs.CVstat.MLarXiv:2009.09796v12020
  54. Image Segmentation Using Text and Image Prompts

    Timo Lüddecke, Alexander S. Ecker

    cs.CVarXiv:2112.10003v22021
  55. AnaDiffusion: Anatomically CompositionalLatent Diffusion for Controllable 3D Brain MRI Generation

    Huiwen Han, Lulin Liu, Bangya Liu +8

    cs.CVarXiv:2608.23014v12026
  56. Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision

    Aadarsh Sahoo, Georgia Gkioxari

    cs.CVarXiv:2602.13195v12026
  57. MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image Segmentation

    Taha Koleilat, Hojat Asgariandehkordi, Omid Nejati Manzari +3

    cs.CVcs.CLarXiv:2602.20423v12026
  58. Light4D: Training-Free Extreme Viewpoint 4D Video Relighting

    Zhenghuang Wu, Kang Chen, Zeyu Zhang +1

    cs.CVarXiv:2602.11769v12026
  59. UniT: Unified Multimodal Chain-of-Thought Test-time Scaling

    Leon Liangyu Chen, Haoyu Ma, Zhipeng Fan +11

    cs.CVcs.AIcs.LGarXiv:2602.12279v22026
  60. What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis

    Xirui Li, Ming Li, Tianyi Zhou

    cs.CVcs.AIarXiv:2602.12395v12026