Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,401 to 17,460 of 18,820
PraNet: Parallel Reverse Attention Network for Polyp Segmentation
Deng-Ping Fan, Ge-Peng Ji, Tao Zhou +4
eess.IVcs.CVarXiv:2006.11392v42020Evaluating Object Hallucination in Large Vision-Language Models
Yifan Li, Yifan Du, Kun Zhou +3
cs.CVcs.CLcs.MMarXiv:2305.10355v32023Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
Jiazheng Xing, Hangjie Yuan, Lingling Cai +9
cs.CVcs.AIarXiv:2605.31603v22026DenseFuse: A Fusion Approach to Infrared and Visible Images
Hui Li, Xiao-Jun Wu
cs.CVarXiv:1804.08361v92018Deep learning in remote sensing: a review
Xiao Xiang Zhu, Devis Tuia, Lichao Mou +4
cs.CVeess.IVarXiv:1710.03959v12017Focal and Efficient IOU Loss for Accurate Bounding Box Regression
Yi-Fan Zhang, Weiqiang Ren, Zhang Zhang +3
cs.CVarXiv:2101.08158v22021Cascade R-CNN: High Quality Object Detection and Instance Segmentation
Zhaowei Cai, Nuno Vasconcelos
cs.CVarXiv:1906.09756v12019SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
Yulu Pan, Han Yi, Seongsu Ha +4
cs.CVarXiv:2605.31529v22026Image Transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit +4
cs.CVarXiv:1802.05751v32018How can embedding models bind concepts?
Arnas Uselis, Darina Koishigarina, Seong Joon Oh
cs.CVcs.LGarXiv:2605.31503v12026Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Hang Zhang, Xin Li, Lidong Bing
cs.CLcs.CVcs.SDarXiv:2306.02858v42023PointConv: Deep Convolutional Networks on 3D Point Clouds
Wenxuan Wu, Zhongang Qi, Li Fuxin
cs.CVarXiv:1811.07246v32018Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney +6
cs.CVcs.AIcs.CLarXiv:1711.07280v32017Learning from Simulated and Unsupervised Images through Adversarial Training
Ashish Shrivastava, Tomas Pfister, Oncel Tuzel +3
cs.CVcs.LGcs.NEarXiv:1612.07828v22016Linear Scaling Video VLMs for Long Video Understanding
Cristobal Eyzaguirre, Jiajun Wu, Juan Carlos Niebles
cs.CVarXiv:2605.31598v12026Count Anything
Mengqi Lei, Shuokun Cheng, Wei Bao +4
cs.CVarXiv:2605.30846v12026DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
Zhenhao Yang, Xiaoshi Wu, Zhengyao Lv +5
cs.CVarXiv:2605.31336v120263D-R2N2: A Unified Approach for Single and Multi-view 3D Object Reconstruction
Christopher B. Choy, Danfei Xu, JunYoung Gwak +2
cs.CVcs.AIarXiv:1604.00449v12016Task-Focused Memorization for Multimodal Agents
Tao Zou, Yichen He, Tian Qiu +2
cs.CVarXiv:2605.31075v12026Function2Scene: 3D Indoor Scene Layout from Functional Specifications
Ruiqi Wang, Qimin Chen, Daniel Ritchie +4
cs.CVcs.GRarXiv:2605.30819v12026NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections
Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi +3
cs.CVcs.GRcs.LGarXiv:2008.02268v32020Understanding Convolution for Semantic Segmentation
Panqu Wang, Pengfei Chen, Ye Yuan +4
cs.CVarXiv:1702.08502v32017StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement
Junwon Seo, Sushant Veer, Ran Tian +6
cs.CVcs.AIcs.LGarXiv:2606.00267v12026Learning Spatially Regularized Correlation Filters for Visual Tracking
Martin Danelljan, Gustav Häger, Fahad Shahbaz Khan +1
cs.CVarXiv:1608.05571v12016Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks
Jierun Chen, Shiu-hong Kao, Hao He +4
cs.CVarXiv:2303.03667v32023ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders
Sanghyun Woo, Shoubhik Debnath, Ronghang Hu +4
cs.CVarXiv:2301.00808v12023Understanding Neural Networks Through Deep Visualization
Jason Yosinski, Jeff Clune, Anh Nguyen +2
cs.CVcs.LGcs.NEarXiv:1506.06579v12015GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
Yue Cao, Jiarui Xu, Stephen Lin +2
cs.CVcs.AIcs.LGarXiv:1904.11492v12019FFA-Net: Feature Fusion Attention Network for Single Image Dehazing
Xu Qin, Zhilin Wang, Yuanchao Bai +2
cs.CVarXiv:1911.07559v22019AutoAugment: Learning Augmentation Policies from Data
Ekin D. Cubuk, Barret Zoph, Dandelion Mane +2
cs.CVcs.LGstat.MLarXiv:1805.09501v32018Segmenter: Transformer for Semantic Segmentation
Robin Strudel, Ricardo Garcia, Ivan Laptev +1
cs.CVcs.AIcs.LGarXiv:2105.05633v32021Kvasir-SEG: A Segmented Polyp Dataset
Debesh Jha, Pia H. Smedsrud, Michael A. Riegler +4
eess.IVcs.CVarXiv:1911.07069v12019Geometric deep learning on graphs and manifolds using mixture model CNNs
Federico Monti, Davide Boscaini, Jonathan Masci +3
cs.CVarXiv:1611.08402v32016CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
Xuehang Guo, Pingyue Zhang, Ruiyi Zhang +6
cs.CVcs.AIcs.CLarXiv:2608.02833v12026AMASS: Archive of Motion Capture as Surface Shapes
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje +2
cs.CVcs.GRarXiv:1904.03278v12019EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing
Jiayi Song, Shijie Huang, Fangtai Wu +5
cs.CVarXiv:2608.18063v12026OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Qiushi Sun, Kanzhi Cheng, Yian Wang +20
cs.AIcs.CLcs.CVarXiv:2607.28609v22026Beacon: Knowing When and How to Perform Agentic Visual Reasoning
Qixun Wang, Yang Shi, Letian Cheng +11
cs.CVarXiv:2607.28595v12026PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
Mingyang Song, Luxin Xu, Haoyu Sun +3
cs.CVcs.AIcs.CLarXiv:2607.05910v12026Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Runmao Yao, Kairui Hu, Yukang Cao +11
cs.CVarXiv:2607.16401v12026MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators
Yushi Huang, Xiangxin Zhou, Jun Zhang +2
cs.CVcs.LGarXiv:2607.15273v22026RefGC-SR$^2$: Reference-guided Super-Resolution and Refinement of AI Generated Content
Jeahun Sung, Dahyeon Kye, Soo Ye Kim +1
cs.CVarXiv:2606.15158v22026MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
Yujie Wei, Yujin Han, Zhekai Chen +20
cs.CVarXiv:2605.20183v42026Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, Yann LeCun
cs.LGcs.CVstat.MLarXiv:1511.05440v62015Ego4D: Around the World in 3,000 Hours of Egocentric Video
Kristen Grauman, Andrew Westbury, Eugene Byrne +82
cs.CVcs.AIarXiv:2110.07058v32021AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
Tao Xu, Pengchuan Zhang, Qiuyuan Huang +4
cs.CVarXiv:1711.10485v12017Person Transfer GAN to Bridge Domain Gap for Person Re-Identification
Longhui Wei, Shiliang Zhang, Wen Gao +1
cs.CVarXiv:1711.08565v22017Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods
Nicholas Carlini, David Wagner
cs.LGcs.CRcs.CVarXiv:1705.07263v22017Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection
Xiang Li, Wenhai Wang, Lijun Wu +5
cs.CVarXiv:2006.04388v12020Natural Adversarial Examples
Dan Hendrycks, Kevin Zhao, Steven Basart +2
cs.LGcs.CVstat.MLarXiv:1907.07174v42019BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers
Zhiqi Li, Wenhai Wang, Hongyang Li +5
cs.CVarXiv:2203.17270v22022Covid-19: Automatic detection from X-Ray images utilizing Transfer Learning with Convolutional Neural Networks
Ioannis D. Apostolopoulos, Tzani Bessiana
eess.IVcs.CVcs.LGarXiv:2003.11617v12020Ensemble deep learning: A review
M. A. Ganaie, Minghui Hu, A. K. Malik +2
cs.LGcs.AIcs.CVarXiv:2104.02395v32021Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
Longlong Jing, Yingli Tian
cs.CVarXiv:1902.06162v12019GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee +1
cs.CVarXiv:1711.02257v42017PaintBench: Deterministic Evaluation of Precise Visual Editing
Kai Xu, Ellis Brown, Shrikar Madhu +3
cs.GRcs.CVcs.LGarXiv:2606.00188v12026Relational Knowledge Distillation
Wonpyo Park, Dongju Kim, Yan Lu +1
cs.CVcs.LGarXiv:1904.05068v22019SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
Tianhui Liu, Jie Feng, Zhiheng Zheng +6
cs.CVcs.AIcs.CLarXiv:2605.31148v12026Stacked Attention Networks for Image Question Answering
Zichao Yang, Xiaodong He, Jianfeng Gao +2
cs.LGcs.CLcs.CVarXiv:1511.02274v22015Deep Mutual Learning
Ying Zhang, Tao Xiang, Timothy M. Hospedales +1
cs.CVarXiv:1706.00384v12017