Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

17,461 to 17,520 of 18,839

  1. RefGC-SR$^2$: Reference-guided Super-Resolution and Refinement of AI Generated Content

    Jeahun Sung, Dahyeon Kye, Soo Ye Kim +1

    cs.CVarXiv:2606.15158v22026
  2. MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

    Yujie Wei, Yujin Han, Zhekai Chen +20

    cs.CVarXiv:2605.20183v42026
  3. Deep multi-scale video prediction beyond mean square error

    Michael Mathieu, Camille Couprie, Yann LeCun

    cs.LGcs.CVstat.MLarXiv:1511.05440v62015
  4. Ego4D: Around the World in 3,000 Hours of Egocentric Video

    Kristen Grauman, Andrew Westbury, Eugene Byrne +82

    cs.CVcs.AIarXiv:2110.07058v32021
  5. AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks

    Tao Xu, Pengchuan Zhang, Qiuyuan Huang +4

    cs.CVarXiv:1711.10485v12017
  6. Person Transfer GAN to Bridge Domain Gap for Person Re-Identification

    Longhui Wei, Shiliang Zhang, Wen Gao +1

    cs.CVarXiv:1711.08565v22017
  7. Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods

    Nicholas Carlini, David Wagner

    cs.LGcs.CRcs.CVarXiv:1705.07263v22017
  8. Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection

    Xiang Li, Wenhai Wang, Lijun Wu +5

    cs.CVarXiv:2006.04388v12020
  9. Natural Adversarial Examples

    Dan Hendrycks, Kevin Zhao, Steven Basart +2

    cs.LGcs.CVstat.MLarXiv:1907.07174v42019
  10. BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers

    Zhiqi Li, Wenhai Wang, Hongyang Li +5

    cs.CVarXiv:2203.17270v22022
  11. Covid-19: Automatic detection from X-Ray images utilizing Transfer Learning with Convolutional Neural Networks

    Ioannis D. Apostolopoulos, Tzani Bessiana

    eess.IVcs.CVcs.LGarXiv:2003.11617v12020
  12. Ensemble deep learning: A review

    M. A. Ganaie, Minghui Hu, A. K. Malik +2

    cs.LGcs.AIcs.CVarXiv:2104.02395v32021
  13. Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey

    Longlong Jing, Yingli Tian

    cs.CVarXiv:1902.06162v12019
  14. GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks

    Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee +1

    cs.CVarXiv:1711.02257v42017
  15. PaintBench: Deterministic Evaluation of Precise Visual Editing

    Kai Xu, Ellis Brown, Shrikar Madhu +3

    cs.GRcs.CVcs.LGarXiv:2606.00188v12026
  16. Relational Knowledge Distillation

    Wonpyo Park, Dongju Kim, Yan Lu +1

    cs.CVcs.LGarXiv:1904.05068v22019
  17. SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

    Tianhui Liu, Jie Feng, Zhiheng Zheng +6

    cs.CVcs.AIcs.CLarXiv:2605.31148v12026
  18. Stacked Attention Networks for Image Question Answering

    Zichao Yang, Xiaodong He, Jianfeng Gao +2

    cs.LGcs.CLcs.CVarXiv:1511.02274v22015
  19. Deep Mutual Learning

    Ying Zhang, Tao Xiang, Timothy M. Hospedales +1

    cs.CVarXiv:1706.00384v12017
  20. DRAW: A Recurrent Neural Network For Image Generation

    Karol Gregor, Ivo Danihelka, Alex Graves +2

    cs.CVcs.LGcs.NEarXiv:1502.04623v22015
  21. Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs

    Xin Gao, Cheng Yang, Chufan Shi +1

    cs.CLcs.CVarXiv:2606.00477v12026
  22. Deeper Depth Prediction with Fully Convolutional Residual Networks

    Iro Laina, Christian Rupprecht, Vasileios Belagiannis +2

    cs.CVarXiv:1606.00373v22016
  23. HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers

    Issa Sugiura, Shuhei Kurita, Yusuke Oda +1

    cs.CVarXiv:2606.01132v12026
  24. Agent Skills Should Go Beyond Text: The Case for Visual Skills

    Binxiao Xu, Ruichuan An, Bocheng Zou +1

    cs.CVarXiv:2606.01414v12026
  25. Deep Bayesian Active Learning with Image Data

    Yarin Gal, Riashat Islam, Zoubin Ghahramani

    cs.LGcs.CVstat.MLarXiv:1703.02910v12017
  26. Identifying the Best Machine Learning Algorithms for Brain Tumor Segmentation, Progression Assessment, and Overall Survival Prediction in the BRATS Challenge

    Spyridon Bakas, Mauricio Reyes, Andras Jakab +424

    cs.CVcs.AIcs.LGarXiv:1811.02629v32018
  27. Habitat: A Platform for Embodied AI Research

    Manolis Savva, Abhishek Kadian, Oleksandr Maksymets +9

    cs.CVcs.AIcs.CLarXiv:1904.01201v22019
  28. Quality-Guided Semi-Supervised Learning for Medical Image Segmentation

    Kumar Abhishek, Ghassan Hamarneh

    cs.CVarXiv:2606.01753v12026
  29. PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps

    Junlin Long, Zeyu Zhang, Xu Deng +5

    cs.CVarXiv:2606.01788v12026
  30. Alias-Free Generative Adversarial Networks

    Tero Karras, Miika Aittala, Samuli Laine +4

    cs.CVcs.AIcs.LGarXiv:2106.12423v42021
    Summaries:한국어
  31. Gradient Surgery for Multi-Task Learning

    Tianhe Yu, Saurabh Kumar, Abhishek Gupta +3

    cs.LGcs.CVcs.ROarXiv:2001.06782v42020
  32. Depth Anything V2

    Lihe Yang, Bingyi Kang, Zilong Huang +4

    cs.CVarXiv:2406.09414v12024
  33. X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding

    Peiwen Sun, Xudong Lu, Huadai Liu +10

    cs.CVarXiv:2606.02482v32026
  34. Free-Form Image Inpainting with Gated Convolution

    Jiahui Yu, Zhe Lin, Jimei Yang +3

    cs.CVcs.GRcs.LGarXiv:1806.03589v22018
  35. D-NeRF: Neural Radiance Fields for Dynamic Scenes

    Albert Pumarola, Enric Corona, Gerard Pons-Moll +1

    cs.CVarXiv:2011.13961v12020
  36. Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data

    Xintao Wang, Liangbin Xie, Chao Dong +1

    eess.IVcs.CVarXiv:2107.10833v22021
  37. Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

    Lihe Yang, Bingyi Kang, Zilong Huang +3

    cs.CVarXiv:2401.10891v22024
  38. DSSD : Deconvolutional Single Shot Detector

    Cheng-Yang Fu, Wei Liu, Ananth Ranga +2

    cs.CVarXiv:1701.06659v12017
  39. AFUN: Towards an Affordance Foundation Model for Functionality Understanding

    Zhaoning Wang, Yi Zhong, Jiawei Fu +2

    cs.ROcs.CVarXiv:2606.02551v12026
  40. Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

    Lianghui Zhu, Bencheng Liao, Qian Zhang +3

    cs.CVcs.LGarXiv:2401.09417v32024
  41. RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds

    Qingyong Hu, Bo Yang, Linhai Xie +5

    cs.CVcs.LGeess.IVarXiv:1911.11236v32019
  42. CE-Net: Context Encoder Network for 2D Medical Image Segmentation

    Zaiwang Gu, Jun Cheng, Huazhu Fu +6

    cs.CVarXiv:1903.02740v12019
  43. Return of Frustratingly Easy Domain Adaptation

    Baochen Sun, Jiashi Feng, Kate Saenko

    cs.CVcs.AIcs.LGarXiv:1511.05547v22015
  44. Simple Baselines for Human Pose Estimation and Tracking

    Bin Xiao, Haiping Wu, Yichen Wei

    cs.CVarXiv:1804.06208v22018
  45. Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting

    Georgios Tsoumplekas, Stella Bounareli, Vasileios Argyriou

    cs.CVcs.LGarXiv:2606.03792v12026
  46. Noise2Noise: Learning Image Restoration without Clean Data

    Jaakko Lehtinen, Jacob Munkberg, Jon Hasselgren +4

    cs.CVcs.LGstat.MLarXiv:1803.04189v32018
  47. Real-world Anomaly Detection in Surveillance Videos

    Waqas Sultani, Chen Chen, Mubarak Shah

    cs.CVarXiv:1801.04264v32018
  48. Per-Pixel Classification is Not All You Need for Semantic Segmentation

    Bowen Cheng, Alexander G. Schwing, Alexander Kirillov

    cs.CVarXiv:2107.06278v22021
  49. StarGAN v2: Diverse Image Synthesis for Multiple Domains

    Yunjey Choi, Youngjung Uh, Jaejun Yoo +1

    cs.CVcs.LGarXiv:1912.01865v22019
  50. Imagen Video: High Definition Video Generation with Diffusion Models

    Jonathan Ho, William Chan, Chitwan Saharia +8

    cs.CVcs.LGarXiv:2210.02303v12022
  51. Learning Deep CNN Denoiser Prior for Image Restoration

    Kai Zhang, Wangmeng Zuo, Shuhang Gu +1

    cs.CVarXiv:1704.03264v12017
  52. Improving Factuality and Reasoning in Language Models through Multiagent Debate

    Yilun Du, Shuang Li, Antonio Torralba +2

    cs.CLcs.AIcs.CVarXiv:2305.14325v12023
  53. Learning to Discover Cross-Domain Relations with Generative Adversarial Networks

    Taeksoo Kim, Moonsu Cha, Hyunsoo Kim +2

    cs.CVarXiv:1703.05192v22017
  54. MultiResUNet : Rethinking the U-Net Architecture for Multimodal Biomedical Image Segmentation

    Nabil Ibtehaz, M. Sohel Rahman

    cs.CVarXiv:1902.04049v12019
  55. An Underwater Image Enhancement Benchmark Dataset and Beyond

    Chongyi Li, Chunle Guo, Wenqi Ren +4

    cs.CVarXiv:1901.05495v22019
  56. SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation

    Junxiao Yang, Minghao Zhang, Xiaoce Wang +4

    cs.CVcs.AIarXiv:2606.03348v12026
  57. BA-T: An Iterative Transformer for Two-View Bundle Adjustment

    Ganlin Zhang, Weirong Chen, Daniel Cremers +1

    cs.CVarXiv:2606.03287v12026
  58. Bootstrap Your Generator: Unpaired Visual Editing with Flow Matching

    Yoad Tewel, Yuval Atzmon, Gal Chechik +1

    cs.CVarXiv:2606.03911v12026
  59. Benchmarking Visual State Tracking in Multimodal Video Understanding

    Sihyun Yu, Nanye Ma, Pinzhi Huang +8

    cs.CVarXiv:2606.03920v12026
  60. DualGAN: Unsupervised Dual Learning for Image-to-Image Translation

    Zili Yi, Hao Zhang, Ping Tan +1

    cs.CVarXiv:1704.02510v42017