Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,461 to 17,520 of 18,839
RefGC-SR$^2$: Reference-guided Super-Resolution and Refinement of AI Generated Content
Jeahun Sung, Dahyeon Kye, Soo Ye Kim +1
cs.CVarXiv:2606.15158v22026MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
Yujie Wei, Yujin Han, Zhekai Chen +20
cs.CVarXiv:2605.20183v42026Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, Yann LeCun
cs.LGcs.CVstat.MLarXiv:1511.05440v62015Ego4D: Around the World in 3,000 Hours of Egocentric Video
Kristen Grauman, Andrew Westbury, Eugene Byrne +82
cs.CVcs.AIarXiv:2110.07058v32021AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
Tao Xu, Pengchuan Zhang, Qiuyuan Huang +4
cs.CVarXiv:1711.10485v12017Person Transfer GAN to Bridge Domain Gap for Person Re-Identification
Longhui Wei, Shiliang Zhang, Wen Gao +1
cs.CVarXiv:1711.08565v22017Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods
Nicholas Carlini, David Wagner
cs.LGcs.CRcs.CVarXiv:1705.07263v22017Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection
Xiang Li, Wenhai Wang, Lijun Wu +5
cs.CVarXiv:2006.04388v12020Natural Adversarial Examples
Dan Hendrycks, Kevin Zhao, Steven Basart +2
cs.LGcs.CVstat.MLarXiv:1907.07174v42019BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers
Zhiqi Li, Wenhai Wang, Hongyang Li +5
cs.CVarXiv:2203.17270v22022Covid-19: Automatic detection from X-Ray images utilizing Transfer Learning with Convolutional Neural Networks
Ioannis D. Apostolopoulos, Tzani Bessiana
eess.IVcs.CVcs.LGarXiv:2003.11617v12020Ensemble deep learning: A review
M. A. Ganaie, Minghui Hu, A. K. Malik +2
cs.LGcs.AIcs.CVarXiv:2104.02395v32021Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
Longlong Jing, Yingli Tian
cs.CVarXiv:1902.06162v12019GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee +1
cs.CVarXiv:1711.02257v42017PaintBench: Deterministic Evaluation of Precise Visual Editing
Kai Xu, Ellis Brown, Shrikar Madhu +3
cs.GRcs.CVcs.LGarXiv:2606.00188v12026Relational Knowledge Distillation
Wonpyo Park, Dongju Kim, Yan Lu +1
cs.CVcs.LGarXiv:1904.05068v22019SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
Tianhui Liu, Jie Feng, Zhiheng Zheng +6
cs.CVcs.AIcs.CLarXiv:2605.31148v12026Stacked Attention Networks for Image Question Answering
Zichao Yang, Xiaodong He, Jianfeng Gao +2
cs.LGcs.CLcs.CVarXiv:1511.02274v22015Deep Mutual Learning
Ying Zhang, Tao Xiang, Timothy M. Hospedales +1
cs.CVarXiv:1706.00384v12017DRAW: A Recurrent Neural Network For Image Generation
Karol Gregor, Ivo Danihelka, Alex Graves +2
cs.CVcs.LGcs.NEarXiv:1502.04623v22015Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
Xin Gao, Cheng Yang, Chufan Shi +1
cs.CLcs.CVarXiv:2606.00477v12026Deeper Depth Prediction with Fully Convolutional Residual Networks
Iro Laina, Christian Rupprecht, Vasileios Belagiannis +2
cs.CVarXiv:1606.00373v22016HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers
Issa Sugiura, Shuhei Kurita, Yusuke Oda +1
cs.CVarXiv:2606.01132v12026Agent Skills Should Go Beyond Text: The Case for Visual Skills
Binxiao Xu, Ruichuan An, Bocheng Zou +1
cs.CVarXiv:2606.01414v12026Deep Bayesian Active Learning with Image Data
Yarin Gal, Riashat Islam, Zoubin Ghahramani
cs.LGcs.CVstat.MLarXiv:1703.02910v12017Identifying the Best Machine Learning Algorithms for Brain Tumor Segmentation, Progression Assessment, and Overall Survival Prediction in the BRATS Challenge
Spyridon Bakas, Mauricio Reyes, Andras Jakab +424
cs.CVcs.AIcs.LGarXiv:1811.02629v32018Habitat: A Platform for Embodied AI Research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets +9
cs.CVcs.AIcs.CLarXiv:1904.01201v22019Quality-Guided Semi-Supervised Learning for Medical Image Segmentation
Kumar Abhishek, Ghassan Hamarneh
cs.CVarXiv:2606.01753v12026PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps
Junlin Long, Zeyu Zhang, Xu Deng +5
cs.CVarXiv:2606.01788v12026Alias-Free Generative Adversarial Networks
Tero Karras, Miika Aittala, Samuli Laine +4
cs.CVcs.AIcs.LGarXiv:2106.12423v42021Summaries:한국어Gradient Surgery for Multi-Task Learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta +3
cs.LGcs.CVcs.ROarXiv:2001.06782v42020Depth Anything V2
Lihe Yang, Bingyi Kang, Zilong Huang +4
cs.CVarXiv:2406.09414v12024X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
Peiwen Sun, Xudong Lu, Huadai Liu +10
cs.CVarXiv:2606.02482v32026Free-Form Image Inpainting with Gated Convolution
Jiahui Yu, Zhe Lin, Jimei Yang +3
cs.CVcs.GRcs.LGarXiv:1806.03589v22018D-NeRF: Neural Radiance Fields for Dynamic Scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll +1
cs.CVarXiv:2011.13961v12020Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data
Xintao Wang, Liangbin Xie, Chao Dong +1
eess.IVcs.CVarXiv:2107.10833v22021Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
Lihe Yang, Bingyi Kang, Zilong Huang +3
cs.CVarXiv:2401.10891v22024DSSD : Deconvolutional Single Shot Detector
Cheng-Yang Fu, Wei Liu, Ananth Ranga +2
cs.CVarXiv:1701.06659v12017AFUN: Towards an Affordance Foundation Model for Functionality Understanding
Zhaoning Wang, Yi Zhong, Jiawei Fu +2
cs.ROcs.CVarXiv:2606.02551v12026Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
Lianghui Zhu, Bencheng Liao, Qian Zhang +3
cs.CVcs.LGarXiv:2401.09417v32024RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds
Qingyong Hu, Bo Yang, Linhai Xie +5
cs.CVcs.LGeess.IVarXiv:1911.11236v32019CE-Net: Context Encoder Network for 2D Medical Image Segmentation
Zaiwang Gu, Jun Cheng, Huazhu Fu +6
cs.CVarXiv:1903.02740v12019Return of Frustratingly Easy Domain Adaptation
Baochen Sun, Jiashi Feng, Kate Saenko
cs.CVcs.AIcs.LGarXiv:1511.05547v22015Simple Baselines for Human Pose Estimation and Tracking
Bin Xiao, Haiping Wu, Yichen Wei
cs.CVarXiv:1804.06208v22018Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting
Georgios Tsoumplekas, Stella Bounareli, Vasileios Argyriou
cs.CVcs.LGarXiv:2606.03792v12026Noise2Noise: Learning Image Restoration without Clean Data
Jaakko Lehtinen, Jacob Munkberg, Jon Hasselgren +4
cs.CVcs.LGstat.MLarXiv:1803.04189v32018Real-world Anomaly Detection in Surveillance Videos
Waqas Sultani, Chen Chen, Mubarak Shah
cs.CVarXiv:1801.04264v32018Per-Pixel Classification is Not All You Need for Semantic Segmentation
Bowen Cheng, Alexander G. Schwing, Alexander Kirillov
cs.CVarXiv:2107.06278v22021StarGAN v2: Diverse Image Synthesis for Multiple Domains
Yunjey Choi, Youngjung Uh, Jaejun Yoo +1
cs.CVcs.LGarXiv:1912.01865v22019Imagen Video: High Definition Video Generation with Diffusion Models
Jonathan Ho, William Chan, Chitwan Saharia +8
cs.CVcs.LGarXiv:2210.02303v12022Learning Deep CNN Denoiser Prior for Image Restoration
Kai Zhang, Wangmeng Zuo, Shuhang Gu +1
cs.CVarXiv:1704.03264v12017Improving Factuality and Reasoning in Language Models through Multiagent Debate
Yilun Du, Shuang Li, Antonio Torralba +2
cs.CLcs.AIcs.CVarXiv:2305.14325v12023Learning to Discover Cross-Domain Relations with Generative Adversarial Networks
Taeksoo Kim, Moonsu Cha, Hyunsoo Kim +2
cs.CVarXiv:1703.05192v22017MultiResUNet : Rethinking the U-Net Architecture for Multimodal Biomedical Image Segmentation
Nabil Ibtehaz, M. Sohel Rahman
cs.CVarXiv:1902.04049v12019An Underwater Image Enhancement Benchmark Dataset and Beyond
Chongyi Li, Chunle Guo, Wenqi Ren +4
cs.CVarXiv:1901.05495v22019SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation
Junxiao Yang, Minghao Zhang, Xiaoce Wang +4
cs.CVcs.AIarXiv:2606.03348v12026BA-T: An Iterative Transformer for Two-View Bundle Adjustment
Ganlin Zhang, Weirong Chen, Daniel Cremers +1
cs.CVarXiv:2606.03287v12026Bootstrap Your Generator: Unpaired Visual Editing with Flow Matching
Yoad Tewel, Yuval Atzmon, Gal Chechik +1
cs.CVarXiv:2606.03911v12026Benchmarking Visual State Tracking in Multimodal Video Understanding
Sihyun Yu, Nanye Ma, Pinzhi Huang +8
cs.CVarXiv:2606.03920v12026DualGAN: Unsupervised Dual Learning for Image-to-Image Translation
Zili Yi, Hao Zhang, Ping Tan +1
cs.CVarXiv:1704.02510v42017