Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,641 to 17,700 of 18,811
Deep Retinex Decomposition for Low-Light Enhancement
Chen Wei, Wenjing Wang, Wenhan Yang +1
cs.CVarXiv:1808.04560v12018BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation
Changqian Yu, Jingbo Wang, Chao Peng +3
cs.CVarXiv:1808.00897v12018Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
Chen Sun, Abhinav Shrivastava, Saurabh Singh +1
cs.CVcs.AIarXiv:1707.02968v22017UNet 3+: A Full-Scale Connected UNet for Medical Image Segmentation
Huimin Huang, Lanfen Lin, Ruofeng Tong +6
eess.IVcs.CVcs.LGarXiv:2004.08790v12020COVID-Net: A Tailored Deep Convolutional Neural Network Design for Detection of COVID-19 Cases from Chest X-Ray Images
Linda Wang, Alexander Wong
eess.IVcs.CVcs.LGarXiv:2003.09871v42020Selective Kernel Networks
Xiang Li, Wenhai Wang, Xiaolin Hu +1
cs.CVarXiv:1903.06586v22019Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
Yukun Zhu, Ryan Kiros, Richard Zemel +4
cs.CVcs.CLarXiv:1506.06724v12015Skin Lesion Analysis Toward Melanoma Detection: A Challenge at the 2017 International Symposium on Biomedical Imaging (ISBI), Hosted by the International Skin Imaging Collaboration (ISIC)
Noel C. F. Codella, David Gutman, M. Emre Celebi +8
cs.CVarXiv:1710.05006v32017EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera
Christen Millerdurai, Shaoxiang Wang, Yaxu Xie +3
cs.CVcs.GRarXiv:2605.12498v12026A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS
Juan Terven, Diana Cordova-Esparza
cs.CVarXiv:2304.00501v72023Deep Laplacian Pyramid Networks for Fast and Accurate Super-Resolution
Wei-Sheng Lai, Jia-Bin Huang, Narendra Ahuja +1
cs.CVarXiv:1704.03915v22017M2Retinexformer: Multi-Modal Retinexformer for Low-Light Image Enhancement
Youssef Aboelwafa, Hicham G. Elmongui, Marwan Torki
cs.CVarXiv:2605.12556v12026TBD-VLA: Temporal Block Diffusion Vision Language Action Model
Sung-Wook Lee, Xuhui Kang, Yen-Ling Kuo
cs.CVcs.ROarXiv:2606.07895v12026VoxCeleb2: Deep Speaker Recognition
Joon Son Chung, Arsha Nagrani, Andrew Zisserman
cs.SDcs.CVeess.ASarXiv:1806.05622v22018Leveraging Morphology for Historical Script Metrological Analysis
Malamatenia Vlachou Efstathiou, Raphaël Baena, Dominique Stutzmann +1
cs.CVarXiv:2606.09446v12026A Survey of the Recent Architectures of Deep Convolutional Neural Networks
Asifullah Khan, Anabia Sohail, Umme Zahoora +1
cs.CVarXiv:1901.06032v72019Convolutional Two-Stream Network Fusion for Video Action Recognition
Christoph Feichtenhofer, Axel Pinz, Andrew Zisserman
cs.CVarXiv:1604.06573v22016Deeply-Recursive Convolutional Network for Image Super-Resolution
Jiwon Kim, Jung Kwon Lee, Kyoung Mu Lee
cs.CVcs.LGarXiv:1511.04491v22015Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning
Haoran Xu, Hongyu Wang, Yifei Gao +4
cs.CVarXiv:2606.09290v12026Generalised Dice overlap as a deep learning loss function for highly unbalanced segmentations
Carole H Sudre, Wenqi Li, Tom Vercauteren +2
cs.CVarXiv:1707.03237v32017Event-based Vision: A Survey
Guillermo Gallego, Tobi Delbruck, Garrick Orchard +8
cs.CVcs.AIcs.LGarXiv:1904.08405v32019Channel Pruning for Accelerating Very Deep Neural Networks
Yihui He, Xiangyu Zhang, Jian Sun
cs.CVarXiv:1707.06168v22017U-TTT: Towards Generalizable PET Image Denoising via Test-Time Training
Zhiwen Yang, Jiayin Li, Hao Lu +3
cs.CVarXiv:2606.11032v22026Video Diffusion Models
Jonathan Ho, Tim Salimans, Alexey Gritsenko +3
cs.CVcs.AIcs.LGarXiv:2204.03458v22022Conditional Image Generation with PixelCNN Decoders
Aaron van den Oord, Nal Kalchbrenner, Oriol Vinyals +3
cs.CVcs.LGarXiv:1606.05328v22016Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields
Jonathan T. Barron, Ben Mildenhall, Dor Verbin +2
cs.CVcs.GRarXiv:2111.12077v32021Self-training with Noisy Student improves ImageNet classification
Qizhe Xie, Minh-Thang Luong, Eduard Hovy +1
cs.LGcs.CVstat.MLarXiv:1911.04252v42019Universal adversarial perturbations
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi +1
cs.CVcs.AIcs.LGarXiv:1610.08401v32016Learning Efficient Convolutional Networks through Network Slimming
Zhuang Liu, Jianguo Li, Zhiqiang Shen +3
cs.CVcs.AIcs.LGarXiv:1708.06519v12017An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Rinon Gal, Yuval Alaluf, Yuval Atzmon +4
cs.CVcs.CLcs.GRarXiv:2208.01618v12022LLaVA-OneVision: Easy Visual Task Transfer
Bo Li, Yuanhan Zhang, Dong Guo +8
cs.CVcs.AIcs.CLarXiv:2408.03326v32024Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields
Jonathan T. Barron, Ben Mildenhall, Matthew Tancik +3
cs.CVcs.GRarXiv:2103.13415v32021Pixel Recurrent Neural Networks
Aaron van den Oord, Nal Kalchbrenner, Koray Kavukcuoglu
cs.CVcs.LGcs.NEarXiv:1601.06759v32016Deep Domain Confusion: Maximizing for Domain Invariance
Eric Tzeng, Judy Hoffman, Ning Zhang +2
cs.CVarXiv:1412.3474v12014NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
Yuheng Huang, Jianlang Chen, Jiayang Song +6
cs.CVcs.AIcs.MMarXiv:2608.13210v12026PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud
Shaoshuai Shi, Xiaogang Wang, Hongsheng Li
cs.CVarXiv:1812.04244v22018Unsupervised Image-to-Image Translation Networks
Ming-Yu Liu, Thomas Breuel, Jan Kautz
cs.CVcs.AIarXiv:1703.00848v62017An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition
Baoguang Shi, Xiang Bai, Cong Yao
cs.CVarXiv:1507.05717v12015Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-Scale Convolutional Architecture
David Eigen, Rob Fergus
cs.CVarXiv:1411.4734v42014SuperGlue: Learning Feature Matching with Graph Neural Networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz +1
cs.CVarXiv:1911.11763v22019Summaries:한국어VMamba: Visual State Space Model
Yue Liu, Yunjie Tian, Yuzhong Zhao +6
cs.CVarXiv:2401.10166v42024Remote Sensing Image Scene Classification: Benchmark and State of the Art
Gong Cheng, Junwei Han, Xiaoqiang Lu
cs.CVarXiv:1703.00121v12017PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume
Deqing Sun, Xiaodong Yang, Ming-Yu Liu +1
cs.CVarXiv:1709.02371v32017Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare +3
cs.CVcs.AIarXiv:2107.07651v22021Prompt-to-Prompt Image Editing with Cross Attention Control
Amir Hertz, Ron Mokady, Jay Tenenbaum +3
cs.CVcs.CLcs.GRarXiv:2208.01626v12022Convolutional Pose Machines
Shih-En Wei, Varun Ramakrishna, Takeo Kanade +1
cs.CVarXiv:1602.00134v42016Unsupervised Learning of Depth and Ego-Motion from Video
Tinghui Zhou, Matthew Brown, Noah Snavely +1
cs.CVarXiv:1704.07813v22017Image-based Recommendations on Styles and Substitutes
Julian McAuley, Christopher Targett, Qinfeng Shi +1
cs.CVcs.IRarXiv:1506.04757v12015StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks
Han Zhang, Tao Xu, Hongsheng Li +4
cs.CVcs.AIstat.MLarXiv:1612.03242v22016UNETR: Transformers for 3D Medical Image Segmentation
Ali Hatamizadeh, Yucheng Tang, Vishwesh Nath +5
eess.IVcs.CVcs.LGarXiv:2103.10504v32021Swin Transformer V2: Scaling Up Capacity and Resolution
Ze Liu, Han Hu, Yutong Lin +9
cs.CVarXiv:2111.09883v22021CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten +3
cs.CVcs.CLcs.LGarXiv:1612.06890v12016Microsoft COCO Captions: Data Collection and Evaluation Server
Xinlei Chen, Hao Fang, Tsung-Yi Lin +4
cs.CVcs.CLarXiv:1504.00325v22015MatConvNet - Convolutional Neural Networks for MATLAB
Andrea Vedaldi, Karel Lenc
cs.CVcs.LGcs.MSarXiv:1412.4564v32014LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Hao Tan, Mohit Bansal
cs.CLcs.CVcs.LGarXiv:1908.07490v32019Direct Sparse Odometry
Jakob Engel, Vladlen Koltun, Daniel Cremers
cs.CVarXiv:1607.02565v22016DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Hao Zhang, Feng Li, Shilong Liu +5
cs.CVarXiv:2203.03605v42022VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
Yinming Huang, Shuyuan Tu, Xi Yan +5
cs.CVarXiv:2608.18607v22026CosFace: Large Margin Cosine Loss for Deep Face Recognition
Hao Wang, Yitong Wang, Zheng Zhou +5
cs.CVarXiv:1801.09414v22018Efficient Neural Architecture Search via Parameter Sharing
Hieu Pham, Melody Y. Guan, Barret Zoph +2
cs.LGcs.CLcs.CVarXiv:1802.03268v22018