Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,701 to 17,760 of 18,855
Convolutional Two-Stream Network Fusion for Video Action Recognition
Christoph Feichtenhofer, Axel Pinz, Andrew Zisserman
cs.CVarXiv:1604.06573v22016Deeply-Recursive Convolutional Network for Image Super-Resolution
Jiwon Kim, Jung Kwon Lee, Kyoung Mu Lee
cs.CVcs.LGarXiv:1511.04491v22015Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning
Haoran Xu, Hongyu Wang, Yifei Gao +4
cs.CVarXiv:2606.09290v12026Generalised Dice overlap as a deep learning loss function for highly unbalanced segmentations
Carole H Sudre, Wenqi Li, Tom Vercauteren +2
cs.CVarXiv:1707.03237v32017Event-based Vision: A Survey
Guillermo Gallego, Tobi Delbruck, Garrick Orchard +8
cs.CVcs.AIcs.LGarXiv:1904.08405v32019Channel Pruning for Accelerating Very Deep Neural Networks
Yihui He, Xiangyu Zhang, Jian Sun
cs.CVarXiv:1707.06168v22017U-TTT: Towards Generalizable PET Image Denoising via Test-Time Training
Zhiwen Yang, Jiayin Li, Hao Lu +3
cs.CVarXiv:2606.11032v22026Video Diffusion Models
Jonathan Ho, Tim Salimans, Alexey Gritsenko +3
cs.CVcs.AIcs.LGarXiv:2204.03458v22022Conditional Image Generation with PixelCNN Decoders
Aaron van den Oord, Nal Kalchbrenner, Oriol Vinyals +3
cs.CVcs.LGarXiv:1606.05328v22016Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields
Jonathan T. Barron, Ben Mildenhall, Dor Verbin +2
cs.CVcs.GRarXiv:2111.12077v32021Self-training with Noisy Student improves ImageNet classification
Qizhe Xie, Minh-Thang Luong, Eduard Hovy +1
cs.LGcs.CVstat.MLarXiv:1911.04252v42019Universal adversarial perturbations
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi +1
cs.CVcs.AIcs.LGarXiv:1610.08401v32016Learning Efficient Convolutional Networks through Network Slimming
Zhuang Liu, Jianguo Li, Zhiqiang Shen +3
cs.CVcs.AIcs.LGarXiv:1708.06519v12017An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Rinon Gal, Yuval Alaluf, Yuval Atzmon +4
cs.CVcs.CLcs.GRarXiv:2208.01618v12022LLaVA-OneVision: Easy Visual Task Transfer
Bo Li, Yuanhan Zhang, Dong Guo +8
cs.CVcs.AIcs.CLarXiv:2408.03326v32024Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields
Jonathan T. Barron, Ben Mildenhall, Matthew Tancik +3
cs.CVcs.GRarXiv:2103.13415v32021Pixel Recurrent Neural Networks
Aaron van den Oord, Nal Kalchbrenner, Koray Kavukcuoglu
cs.CVcs.LGcs.NEarXiv:1601.06759v32016Deep Domain Confusion: Maximizing for Domain Invariance
Eric Tzeng, Judy Hoffman, Ning Zhang +2
cs.CVarXiv:1412.3474v12014NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
Yuheng Huang, Jianlang Chen, Jiayang Song +6
cs.CVcs.AIcs.MMarXiv:2608.13210v12026PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud
Shaoshuai Shi, Xiaogang Wang, Hongsheng Li
cs.CVarXiv:1812.04244v22018Unsupervised Image-to-Image Translation Networks
Ming-Yu Liu, Thomas Breuel, Jan Kautz
cs.CVcs.AIarXiv:1703.00848v62017An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition
Baoguang Shi, Xiang Bai, Cong Yao
cs.CVarXiv:1507.05717v12015Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-Scale Convolutional Architecture
David Eigen, Rob Fergus
cs.CVarXiv:1411.4734v42014SuperGlue: Learning Feature Matching with Graph Neural Networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz +1
cs.CVarXiv:1911.11763v22019Summaries:한국어VMamba: Visual State Space Model
Yue Liu, Yunjie Tian, Yuzhong Zhao +6
cs.CVarXiv:2401.10166v42024Remote Sensing Image Scene Classification: Benchmark and State of the Art
Gong Cheng, Junwei Han, Xiaoqiang Lu
cs.CVarXiv:1703.00121v12017PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume
Deqing Sun, Xiaodong Yang, Ming-Yu Liu +1
cs.CVarXiv:1709.02371v32017Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare +3
cs.CVcs.AIarXiv:2107.07651v22021Prompt-to-Prompt Image Editing with Cross Attention Control
Amir Hertz, Ron Mokady, Jay Tenenbaum +3
cs.CVcs.CLcs.GRarXiv:2208.01626v12022Convolutional Pose Machines
Shih-En Wei, Varun Ramakrishna, Takeo Kanade +1
cs.CVarXiv:1602.00134v42016Unsupervised Learning of Depth and Ego-Motion from Video
Tinghui Zhou, Matthew Brown, Noah Snavely +1
cs.CVarXiv:1704.07813v22017Image-based Recommendations on Styles and Substitutes
Julian McAuley, Christopher Targett, Qinfeng Shi +1
cs.CVcs.IRarXiv:1506.04757v12015StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks
Han Zhang, Tao Xu, Hongsheng Li +4
cs.CVcs.AIstat.MLarXiv:1612.03242v22016UNETR: Transformers for 3D Medical Image Segmentation
Ali Hatamizadeh, Yucheng Tang, Vishwesh Nath +5
eess.IVcs.CVcs.LGarXiv:2103.10504v32021Swin Transformer V2: Scaling Up Capacity and Resolution
Ze Liu, Han Hu, Yutong Lin +9
cs.CVarXiv:2111.09883v22021CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten +3
cs.CVcs.CLcs.LGarXiv:1612.06890v12016Microsoft COCO Captions: Data Collection and Evaluation Server
Xinlei Chen, Hao Fang, Tsung-Yi Lin +4
cs.CVcs.CLarXiv:1504.00325v22015MatConvNet - Convolutional Neural Networks for MATLAB
Andrea Vedaldi, Karel Lenc
cs.CVcs.LGcs.MSarXiv:1412.4564v32014LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Hao Tan, Mohit Bansal
cs.CLcs.CVcs.LGarXiv:1908.07490v32019Direct Sparse Odometry
Jakob Engel, Vladlen Koltun, Daniel Cremers
cs.CVarXiv:1607.02565v22016DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Hao Zhang, Feng Li, Shilong Liu +5
cs.CVarXiv:2203.03605v42022VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
Yinming Huang, Shuyuan Tu, Xi Yan +5
cs.CVarXiv:2608.18607v22026CosFace: Large Margin Cosine Loss for Deep Face Recognition
Hao Wang, Yitong Wang, Zheng Zhou +5
cs.CVarXiv:1801.09414v22018Efficient Neural Architecture Search via Parameter Sharing
Hieu Pham, Melody Y. Guan, Barret Zoph +2
cs.LGcs.CLcs.CVarXiv:1802.03268v22018Towards Real-Time and Adaptable LiDAR Scene Completion
Azhar Hussian, Martin Vossiek, Vasileios Belagiannis
cs.CVarXiv:2608.16490v12026SPK: Eliciting Structured Prior Knowledge for Interpretable Out-of-Distribution Detection in Real-Time Object Detection
Changshun Wu, Weicheng He, Xiaowei Huang +1
cs.CVcs.LGarXiv:2608.19080v12026Unsupervised Visual Representation Learning by Context Prediction
Carl Doersch, Abhinav Gupta, Alexei A. Efros
cs.CVarXiv:1505.05192v32015Cyclical Learning Rates for Training Neural Networks
Leslie N. Smith
cs.CVcs.LGcs.NEarXiv:1506.01186v62015DOTA: A Large-scale Dataset for Object Detection in Aerial Images
Gui-Song Xia, Xiang Bai, Jian Ding +6
cs.CVarXiv:1711.10398v32017NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis
Amir Shahroudy, Jun Liu, Tian-Tsong Ng +1
cs.CVarXiv:1604.02808v12016A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation
Nikolaus Mayer, Eddy Ilg, Philip Häusser +4
cs.CVcs.LGstat.MLarXiv:1512.02134v12015Fine-Grained Visual Classification of Aircraft
Subhransu Maji, Esa Rahtu, Juho Kannala +2
cs.CVarXiv:1306.5151v12013WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
Hengyuan Xu, Qixun Wang, Yiji Cheng +5
cs.CVarXiv:2608.20336v12026CCNet: Criss-Cross Attention for Semantic Segmentation
Zilong Huang, Xinggang Wang, Yunchao Wei +4
cs.CVarXiv:1811.11721v22018Self-Evolving Visual Questioner
Yijun Liang, Hengguang Zhou, Ming Li +3
cs.CVcs.LGarXiv:2606.13929v12026Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering
Mohammed Abdul Al Arafat Tanzin, Rudzidatul Akmam Dziyauddin
cs.CVcs.AIcs.LGarXiv:2608.14724v12026MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Andrew G. Howard, Menglong Zhu, Bo Chen +5
cs.CVarXiv:1704.04861v12017Attention U-Net: Learning Where to Look for the Pancreas
Ozan Oktay, Jo Schlemper, Loic Le Folgoc +9
cs.CVarXiv:1804.03999v32018LUNG-KGMM: Knowledge-Guided Multimodal Learning for Lung Cancer Incidence Prediction
Chunlei Yang, Shuyan Li, Zhong Cao
cs.LGcs.CVarXiv:2608.14657v12026Summaries:한국어Xception: Deep Learning with Depthwise Separable Convolutions
François Chollet
cs.CVarXiv:1610.02357v32016