Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
15,421 to 15,480 of 18,821
AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
Xilin Jiang, Qiaolin Wang, Junkai Wu +30
cs.SDcs.CLcs.CVarXiv:2601.17645v12026SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding
Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum +2
cs.AIcs.CVarXiv:2601.21666v22026Videos as Space-Time Region Graphs
Xiaolong Wang, Abhinav Gupta
cs.CVarXiv:1806.01810v22018DeepIM: Deep Iterative Matching for 6D Pose Estimation
Yi Li, Gu Wang, Xiangyang Ji +2
cs.CVcs.ROarXiv:1804.00175v42018End-to-End Learning of Visual Representations from Uncurated Instructional Videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira +3
cs.CVarXiv:1912.06430v42019HyperAlign: Hypernetwork for Efficient Test-Time Alignment of Diffusion Models
Xin Xie, Jiaxian Guo, Dong Gong
cs.CVarXiv:2601.15968v22026Embodied Question Answering
Abhishek Das, Samyak Datta, Georgia Gkioxari +3
cs.CVcs.AIcs.CLarXiv:1711.11543v22017DeFM: Learning Foundation Representations from Depth for Robotics
Manthan Patel, Jonas Frey, Mayank Mittal +5
cs.ROcs.CVarXiv:2601.18923v12026Quality Inspection of Printed Circuit Board Pin Insertion via Semantic Segmentation and Board-Level Feature Extraction
Nils Rabeneck, André Kiunke, Nicole Hoess +1
cs.CVarXiv:2608.22937v12026Novel methods for multilinear data completion and de-noising based on tensor-SVD
Zemin Zhang, Gregory Ely, Shuchin Aeron +2
cs.CVarXiv:1407.1785v22014Towards Pixel-Level VLM Perception via Simple Points Prediction
Tianhui Song, Haoyu Lu, Hao Yang +8
cs.CVarXiv:2601.19228v12026Hyperspherical Autoencoder for High-Fidelity Image Reconstruction and Generation
Hun Chang, Byunghee Cha, Jong Chul Ye
cs.CVcs.AIcs.LGarXiv:2601.22904v22026Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning
Yu Xu, Yuxin Zhang, Juan Cao +5
cs.CVcs.AIarXiv:2602.01335v12026DRIT++: Diverse Image-to-Image Translation via Disentangled Representations
Hsin-Ying Lee, Hung-Yu Tseng, Qi Mao +4
cs.CVarXiv:1905.01270v22019SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-body Manipulation
Mu Huang, Hui Wang, Kerui Ren +5
cs.ROcs.AIcs.CVarXiv:2602.02402v22026Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers
Bozhou Li, Yushuo Guan, Haolin Li +7
cs.CVarXiv:2602.03510v12026Delving Deeper into Convolutional Networks for Learning Video Representations
Nicolas Ballas, Li Yao, Chris Pal +1
cs.CVcs.LGcs.NEarXiv:1511.06432v42015Capsule-Forensics: Using Capsule Networks to Detect Forged Images and Videos
Huy H. Nguyen, Junichi Yamagishi, Isao Echizen
cs.CVeess.IVarXiv:1810.11215v12018CGNet: A Light-weight Context Guided Network for Semantic Segmentation
Tianyi Wu, Sheng Tang, Rui Zhang +1
cs.CVarXiv:1811.08201v22018Scaling Up Your Kernels to 31x31: Revisiting Large Kernel Design in CNNs
Xiaohan Ding, Xiangyu Zhang, Yizhuang Zhou +3
cs.CVcs.AIcs.LGarXiv:2203.06717v42022TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving
Kashyap Chitta, Aditya Prakash, Bernhard Jaeger +3
cs.CVcs.AIcs.LGarXiv:2205.15997v12022Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks
Bohan Zeng, Kaixin Zhu, Daili Hua +24
cs.CVarXiv:2602.01630v12026Working hard to know your neighbor's margins: Local descriptor learning loss
Anastasiya Mishchuk, Dmytro Mishkin, Filip Radenovic +1
cs.CVarXiv:1705.10872v420173D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
Zhixue Fang, Xu He, Songlin Tang +5
cs.CVarXiv:2602.03796v22026Attribute2Image: Conditional Image Generation from Visual Attributes
Xinchen Yan, Jimei Yang, Kihyuk Sohn +1
cs.LGcs.AIcs.CVarXiv:1512.00570v22015How Merge-Tolerant Are Vision Transformers for Wheat Phenotyping?
Simon Ravé, Pejman Rasti, David Rousseau
cs.CVarXiv:2608.23142v12026TransTrack: Multiple Object Tracking with Transformer
Peize Sun, Jinkun Cao, Yi Jiang +5
cs.CVarXiv:2012.15460v22020Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields
Dor Verbin, Peter Hedman, Ben Mildenhall +3
cs.CVcs.GRarXiv:2112.03907v12021Adaptive 1D Video Diffusion Autoencoder
Yao Teng, Minxuan Lin, Xian Liu +3
cs.CVarXiv:2602.04220v12026When the Edit Changes the Patient: Measuring Identity Preservation in Counterfactual Retinal Images
Andrea Posada, Wenke Karbole, Bach Ngoc Doan +9
cs.CVarXiv:2608.23024v12026OmniRad: A Radiological Foundation Model for Multi-Task Medical Image Analysis
Luca Zedda, Andrea Loddo, Cecilia Di Ruberto
cs.CVcs.AIarXiv:2602.04547v12026No More Strided Convolutions or Pooling: A New CNN Building Block for Low-Resolution Images and Small Objects
Raja Sunkara, Tie Luo
cs.CVcs.LGarXiv:2208.03641v12022Cross Attention Network for Few-shot Classification
Ruibing Hou, Hong Chang, Bingpeng Ma +2
cs.CVarXiv:1910.07677v12019Human Action Recognition from Various Data Modalities: A Review
Zehua Sun, Qiuhong Ke, Hossein Rahmani +3
cs.CVarXiv:2012.11866v52020When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
Jiacheng Hou, Yining Sun, Ruochong Jin +4
cs.CVcs.AIarXiv:2602.10179v22026Object Goal Navigation using Goal-Oriented Semantic Exploration
Devendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta +1
cs.CVcs.LGcs.ROarXiv:2007.00643v22020Neighbor-Aware View Synthesis for Restoring Missing Views in Light-Field Camera Arrays
Sakshi Goel, Ayush Goyal, K S Venkatesh +1
cs.CVarXiv:2608.23175v12026TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution
Deyang Jiang, Jing Huang, Xuanle Zhao +6
cs.CVarXiv:2602.09662v12026DeepFakes: a New Threat to Face Recognition? Assessment and Detection
Pavel Korshunov, Sebastien Marcel
cs.CVarXiv:1812.08685v12018Reinforced Attention Learning
Bangzheng Li, Jianmo Ni, Chen Qu +5
cs.CLcs.CVcs.LGarXiv:2602.04884v22026Auto-Encoding Scene Graphs for Image Captioning
Xu Yang, Kaihua Tang, Hanwang Zhang +1
cs.CVarXiv:1812.02378v32018Fast-SAM3D: 3Dfy Anything in Images but Faster
Weilun Feng, Mingqiang Wu, Zhiliang Chen +10
cs.CVarXiv:2602.05293v22026Prism: Spectral-Aware Block-Sparse Attention
Xinghao Wang, Pengyu Wang, Xiaoran Liu +4
cs.CLcs.AIcs.CVarXiv:2602.08426v22026MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE
Ruijie Zhu, Jiahao Lu, Wenbo Hu +4
cs.CVcs.AIcs.CGarXiv:2602.08961v22026ArcFlow: Unleashing 2-Step Text-to-Image Generation via High-Precision Non-Linear Flow Distillation
Zihan Yang, Shuyuan Tu, Licheng Zhang +3
cs.CVcs.AIarXiv:2602.09014v12026Joint Optimization Framework for Learning with Noisy Labels
Daiki Tanaka, Daiki Ikami, Toshihiko Yamasaki +1
cs.CVcs.LGstat.MLarXiv:1803.11364v12018Reasoning-Augmented Representations for Multimodal Retrieval
Jianrui Zhang, Anirudh Sundara Rajan, Brandon Han +3
cs.IRcs.AIcs.CVarXiv:2602.07125v12026VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
Issar Tzachor, Dvir Samuel, Rami Ben-Ari
cs.CVcs.AIarXiv:2602.08099v12026Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning
Xu Ma, Yitian Zhang, Qihua Dong +1
cs.CVarXiv:2602.09439v12026Motion-Based Tokenization for Cross-Dataset Egocentric Gaze Modeling
Virmarie Maquiling, Zhuojiang Cai, Enkelejda Kasneci
cs.CVarXiv:2608.22926v12026Video Frame Synthesis using Deep Voxel Flow
Ziwei Liu, Raymond A. Yeh, Xiaoou Tang +2
cs.CVcs.GRcs.LGarXiv:1702.02463v22017Gated Siamese Convolutional Neural Network Architecture for Human Re-Identification
Rahul Rama Varior, Mrinal Haloi, Gang Wang
cs.CVarXiv:1607.08378v22016Multi-Task Learning with Deep Neural Networks: A Survey
Michael Crawshaw
cs.LGcs.CVstat.MLarXiv:2009.09796v12020Image Segmentation Using Text and Image Prompts
Timo Lüddecke, Alexander S. Ecker
cs.CVarXiv:2112.10003v22021AnaDiffusion: Anatomically CompositionalLatent Diffusion for Controllable 3D Brain MRI Generation
Huiwen Han, Lulin Liu, Bangya Liu +8
cs.CVarXiv:2608.23014v12026Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision
Aadarsh Sahoo, Georgia Gkioxari
cs.CVarXiv:2602.13195v12026MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image Segmentation
Taha Koleilat, Hojat Asgariandehkordi, Omid Nejati Manzari +3
cs.CVcs.CLarXiv:2602.20423v12026Light4D: Training-Free Extreme Viewpoint 4D Video Relighting
Zhenghuang Wu, Kang Chen, Zeyu Zhang +1
cs.CVarXiv:2602.11769v12026UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
Leon Liangyu Chen, Haoyu Ma, Zhipeng Fan +11
cs.CVcs.AIcs.LGarXiv:2602.12279v22026What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis
Xirui Li, Ming Li, Tianyi Zhou
cs.CVcs.AIarXiv:2602.12395v12026