Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

17,641 to 17,700 of 18,811

  1. Deep Retinex Decomposition for Low-Light Enhancement

    Chen Wei, Wenjing Wang, Wenhan Yang +1

    cs.CVarXiv:1808.04560v12018
  2. BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation

    Changqian Yu, Jingbo Wang, Chao Peng +3

    cs.CVarXiv:1808.00897v12018
  3. Revisiting Unreasonable Effectiveness of Data in Deep Learning Era

    Chen Sun, Abhinav Shrivastava, Saurabh Singh +1

    cs.CVcs.AIarXiv:1707.02968v22017
  4. UNet 3+: A Full-Scale Connected UNet for Medical Image Segmentation

    Huimin Huang, Lanfen Lin, Ruofeng Tong +6

    eess.IVcs.CVcs.LGarXiv:2004.08790v12020
  5. COVID-Net: A Tailored Deep Convolutional Neural Network Design for Detection of COVID-19 Cases from Chest X-Ray Images

    Linda Wang, Alexander Wong

    eess.IVcs.CVcs.LGarXiv:2003.09871v42020
  6. Selective Kernel Networks

    Xiang Li, Wenhai Wang, Xiaolin Hu +1

    cs.CVarXiv:1903.06586v22019
  7. Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books

    Yukun Zhu, Ryan Kiros, Richard Zemel +4

    cs.CVcs.CLarXiv:1506.06724v12015
  8. Skin Lesion Analysis Toward Melanoma Detection: A Challenge at the 2017 International Symposium on Biomedical Imaging (ISBI), Hosted by the International Skin Imaging Collaboration (ISIC)

    Noel C. F. Codella, David Gutman, M. Emre Celebi +8

    cs.CVarXiv:1710.05006v32017
  9. EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera

    Christen Millerdurai, Shaoxiang Wang, Yaxu Xie +3

    cs.CVcs.GRarXiv:2605.12498v12026
  10. A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS

    Juan Terven, Diana Cordova-Esparza

    cs.CVarXiv:2304.00501v72023
  11. Deep Laplacian Pyramid Networks for Fast and Accurate Super-Resolution

    Wei-Sheng Lai, Jia-Bin Huang, Narendra Ahuja +1

    cs.CVarXiv:1704.03915v22017
  12. M2Retinexformer: Multi-Modal Retinexformer for Low-Light Image Enhancement

    Youssef Aboelwafa, Hicham G. Elmongui, Marwan Torki

    cs.CVarXiv:2605.12556v12026
  13. TBD-VLA: Temporal Block Diffusion Vision Language Action Model

    Sung-Wook Lee, Xuhui Kang, Yen-Ling Kuo

    cs.CVcs.ROarXiv:2606.07895v12026
  14. VoxCeleb2: Deep Speaker Recognition

    Joon Son Chung, Arsha Nagrani, Andrew Zisserman

    cs.SDcs.CVeess.ASarXiv:1806.05622v22018
  15. Leveraging Morphology for Historical Script Metrological Analysis

    Malamatenia Vlachou Efstathiou, Raphaël Baena, Dominique Stutzmann +1

    cs.CVarXiv:2606.09446v12026
  16. A Survey of the Recent Architectures of Deep Convolutional Neural Networks

    Asifullah Khan, Anabia Sohail, Umme Zahoora +1

    cs.CVarXiv:1901.06032v72019
  17. Convolutional Two-Stream Network Fusion for Video Action Recognition

    Christoph Feichtenhofer, Axel Pinz, Andrew Zisserman

    cs.CVarXiv:1604.06573v22016
  18. Deeply-Recursive Convolutional Network for Image Super-Resolution

    Jiwon Kim, Jung Kwon Lee, Kyoung Mu Lee

    cs.CVcs.LGarXiv:1511.04491v22015
  19. Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning

    Haoran Xu, Hongyu Wang, Yifei Gao +4

    cs.CVarXiv:2606.09290v12026
  20. Generalised Dice overlap as a deep learning loss function for highly unbalanced segmentations

    Carole H Sudre, Wenqi Li, Tom Vercauteren +2

    cs.CVarXiv:1707.03237v32017
  21. Event-based Vision: A Survey

    Guillermo Gallego, Tobi Delbruck, Garrick Orchard +8

    cs.CVcs.AIcs.LGarXiv:1904.08405v32019
  22. Channel Pruning for Accelerating Very Deep Neural Networks

    Yihui He, Xiangyu Zhang, Jian Sun

    cs.CVarXiv:1707.06168v22017
  23. U-TTT: Towards Generalizable PET Image Denoising via Test-Time Training

    Zhiwen Yang, Jiayin Li, Hao Lu +3

    cs.CVarXiv:2606.11032v22026
  24. Video Diffusion Models

    Jonathan Ho, Tim Salimans, Alexey Gritsenko +3

    cs.CVcs.AIcs.LGarXiv:2204.03458v22022
  25. Conditional Image Generation with PixelCNN Decoders

    Aaron van den Oord, Nal Kalchbrenner, Oriol Vinyals +3

    cs.CVcs.LGarXiv:1606.05328v22016
  26. Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields

    Jonathan T. Barron, Ben Mildenhall, Dor Verbin +2

    cs.CVcs.GRarXiv:2111.12077v32021
  27. Self-training with Noisy Student improves ImageNet classification

    Qizhe Xie, Minh-Thang Luong, Eduard Hovy +1

    cs.LGcs.CVstat.MLarXiv:1911.04252v42019
  28. Universal adversarial perturbations

    Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi +1

    cs.CVcs.AIcs.LGarXiv:1610.08401v32016
  29. Learning Efficient Convolutional Networks through Network Slimming

    Zhuang Liu, Jianguo Li, Zhiqiang Shen +3

    cs.CVcs.AIcs.LGarXiv:1708.06519v12017
  30. An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

    Rinon Gal, Yuval Alaluf, Yuval Atzmon +4

    cs.CVcs.CLcs.GRarXiv:2208.01618v12022
  31. LLaVA-OneVision: Easy Visual Task Transfer

    Bo Li, Yuanhan Zhang, Dong Guo +8

    cs.CVcs.AIcs.CLarXiv:2408.03326v32024
  32. Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields

    Jonathan T. Barron, Ben Mildenhall, Matthew Tancik +3

    cs.CVcs.GRarXiv:2103.13415v32021
  33. Pixel Recurrent Neural Networks

    Aaron van den Oord, Nal Kalchbrenner, Koray Kavukcuoglu

    cs.CVcs.LGcs.NEarXiv:1601.06759v32016
  34. Deep Domain Confusion: Maximizing for Domain Invariance

    Eric Tzeng, Judy Hoffman, Ning Zhang +2

    cs.CVarXiv:1412.3474v12014
  35. NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

    Yuheng Huang, Jianlang Chen, Jiayang Song +6

    cs.CVcs.AIcs.MMarXiv:2608.13210v12026
  36. PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud

    Shaoshuai Shi, Xiaogang Wang, Hongsheng Li

    cs.CVarXiv:1812.04244v22018
  37. Unsupervised Image-to-Image Translation Networks

    Ming-Yu Liu, Thomas Breuel, Jan Kautz

    cs.CVcs.AIarXiv:1703.00848v62017
  38. An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition

    Baoguang Shi, Xiang Bai, Cong Yao

    cs.CVarXiv:1507.05717v12015
  39. Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-Scale Convolutional Architecture

    David Eigen, Rob Fergus

    cs.CVarXiv:1411.4734v42014
  40. SuperGlue: Learning Feature Matching with Graph Neural Networks

    Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz +1

    cs.CVarXiv:1911.11763v22019
    Summaries:한국어
  41. VMamba: Visual State Space Model

    Yue Liu, Yunjie Tian, Yuzhong Zhao +6

    cs.CVarXiv:2401.10166v42024
  42. Remote Sensing Image Scene Classification: Benchmark and State of the Art

    Gong Cheng, Junwei Han, Xiaoqiang Lu

    cs.CVarXiv:1703.00121v12017
  43. PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume

    Deqing Sun, Xiaodong Yang, Ming-Yu Liu +1

    cs.CVarXiv:1709.02371v32017
  44. Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

    Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare +3

    cs.CVcs.AIarXiv:2107.07651v22021
  45. Prompt-to-Prompt Image Editing with Cross Attention Control

    Amir Hertz, Ron Mokady, Jay Tenenbaum +3

    cs.CVcs.CLcs.GRarXiv:2208.01626v12022
  46. Convolutional Pose Machines

    Shih-En Wei, Varun Ramakrishna, Takeo Kanade +1

    cs.CVarXiv:1602.00134v42016
  47. Unsupervised Learning of Depth and Ego-Motion from Video

    Tinghui Zhou, Matthew Brown, Noah Snavely +1

    cs.CVarXiv:1704.07813v22017
  48. Image-based Recommendations on Styles and Substitutes

    Julian McAuley, Christopher Targett, Qinfeng Shi +1

    cs.CVcs.IRarXiv:1506.04757v12015
  49. StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks

    Han Zhang, Tao Xu, Hongsheng Li +4

    cs.CVcs.AIstat.MLarXiv:1612.03242v22016
  50. UNETR: Transformers for 3D Medical Image Segmentation

    Ali Hatamizadeh, Yucheng Tang, Vishwesh Nath +5

    eess.IVcs.CVcs.LGarXiv:2103.10504v32021
  51. Swin Transformer V2: Scaling Up Capacity and Resolution

    Ze Liu, Han Hu, Yutong Lin +9

    cs.CVarXiv:2111.09883v22021
  52. CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

    Justin Johnson, Bharath Hariharan, Laurens van der Maaten +3

    cs.CVcs.CLcs.LGarXiv:1612.06890v12016
  53. Microsoft COCO Captions: Data Collection and Evaluation Server

    Xinlei Chen, Hao Fang, Tsung-Yi Lin +4

    cs.CVcs.CLarXiv:1504.00325v22015
  54. MatConvNet - Convolutional Neural Networks for MATLAB

    Andrea Vedaldi, Karel Lenc

    cs.CVcs.LGcs.MSarXiv:1412.4564v32014
  55. LXMERT: Learning Cross-Modality Encoder Representations from Transformers

    Hao Tan, Mohit Bansal

    cs.CLcs.CVcs.LGarXiv:1908.07490v32019
  56. Direct Sparse Odometry

    Jakob Engel, Vladlen Koltun, Daniel Cremers

    cs.CVarXiv:1607.02565v22016
  57. DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

    Hao Zhang, Feng Li, Shilong Liu +5

    cs.CVarXiv:2203.03605v42022
  58. VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

    Yinming Huang, Shuyuan Tu, Xi Yan +5

    cs.CVarXiv:2608.18607v22026
  59. CosFace: Large Margin Cosine Loss for Deep Face Recognition

    Hao Wang, Yitong Wang, Zheng Zhou +5

    cs.CVarXiv:1801.09414v22018
  60. Efficient Neural Architecture Search via Parameter Sharing

    Hieu Pham, Melody Y. Guan, Barret Zoph +2

    cs.LGcs.CLcs.CVarXiv:1802.03268v22018