Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
10,141 to 10,200 of 18,867
Fast Alternating Linearization Methods for Minimizing the Sum of Two Convex Functions
Donald Goldfarb, Shiqian Ma, Katya Scheinberg
math.OCcs.CVmath.NAarXiv:0912.4571v22009Beyond Temporal Pooling: Recurrence and Temporal Convolutions for Gesture Recognition in Video
Lionel Pigou, Aäron van den Oord, Sander Dieleman +2
cs.CVcs.AIcs.LGarXiv:1506.01911v32015Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang +2
cs.CVarXiv:2608.28192v12026Monte Carlo Convolution for Learning on Non-Uniformly Sampled Point Clouds
Pedro Hermosilla, Tobias Ritschel, Pere-Pau Vázquez +2
cs.CVarXiv:1806.01759v22018Density Map Guided Object Detection in Aerial Images
Changlin Li, Taojiannan Yang, Sijie Zhu +2
cs.CVarXiv:2004.05520v12020Cross-domain Detection via Graph-induced Prototype Alignment
Minghao Xu, Hang Wang, Bingbing Ni +2
cs.CVarXiv:2003.12849v12020SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation
Wenchao Sun, Xuewu Lin, Yining Shi +3
cs.CVarXiv:2405.19620v22024Person Re-identification by Contour Sketch under Moderate Clothing Change
Qize Yang, Ancong Wu, Wei-Shi Zheng
cs.CVarXiv:2002.02295v12020Cut and Learn for Unsupervised Object Detection and Instance Segmentation
Xudong Wang, Rohit Girdhar, Stella X. Yu +1
cs.CVcs.AIcs.LGarXiv:2301.11320v120233D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans
Antoni Rosinol, Arjun Gupta, Marcus Abate +2
cs.ROcs.AIcs.CVarXiv:2002.06289v22020Uncertainty-aware multi-view co-training for semi-supervised medical image segmentation and domain adaptation
Yingda Xia, Dong Yang, Zhiding Yu +7
cs.CVarXiv:2006.16806v12020GP-GAN: Towards Realistic High-Resolution Image Blending
Huikai Wu, Shuai Zheng, Junge Zhang +1
cs.CVarXiv:1703.07195v32017A Survey of Vision-Language Pre-Trained Models
Yifan Du, Zikang Liu, Junyi Li +1
cs.CVcs.CLcs.LGarXiv:2202.10936v22022Automatic 3D liver location and segmentation via convolutional neural networks and graph cut
Fang Lu, Fa Wu, Peijun Hu +2
cs.CVarXiv:1605.03012v12016GS-IR: 3D Gaussian Splatting for Inverse Rendering
Zhihao Liang, Qi Zhang, Ying Feng +2
cs.CVarXiv:2311.16473v32023Improving Vision-and-Language Navigation with Image-Text Pairs from the Web
Arjun Majumdar, Ayush Shrivastava, Stefan Lee +3
cs.CVcs.AIcs.CLarXiv:2004.14973v22020Beyond the Pixel-Wise Loss for Topology-Aware Delineation
Agata Mosinska, Pablo Marquez-Neila, Mateusz Kozinski +1
cs.CVarXiv:1712.02190v12017Robust and Generalizable Visual Representation Learning via Random Convolutions
Zhenlin Xu, Deyi Liu, Junlin Yang +2
cs.CVcs.LGarXiv:2007.13003v32020Learning a Deep ConvNet for Multi-label Classification with Partial Labels
Thibaut Durand, Nazanin Mehrasa, Greg Mori
cs.CVarXiv:1902.09720v12019LAPAR: Linearly-Assembled Pixel-Adaptive Regression Network for Single Image Super-Resolution and Beyond
Wenbo Li, Kun Zhou, Lu Qi +3
cs.CVarXiv:2105.10422v12021Exploring Spatial Context for 3D Semantic Segmentation of Point Clouds
Francis Engelmann, Theodora Kontogianni, Alexander Hermans +1
cs.CVarXiv:1802.01500v22018LMSCNet: Lightweight Multiscale 3D Semantic Completion
Luis Roldão, Raoul de Charette, Anne Verroust-Blondet
cs.CVarXiv:2008.10559v22020Scene as Occupancy
Chonghao Sima, Wenwen Tong, Tai Wang +8
cs.CVcs.ROarXiv:2306.02851v32023ReNet: A Recurrent Neural Network Based Alternative to Convolutional Networks
Francesco Visin, Kyle Kastner, Kyunghyun Cho +3
cs.CVarXiv:1505.00393v32015Guided Motion Diffusion for Controllable Human Motion Synthesis
Korrawe Karunratanakul, Konpat Preechakul, Supasorn Suwajanakorn +1
cs.CVarXiv:2305.12577v32023CSPN++: Learning Context and Resource Aware Convolutional Spatial Propagation Networks for Depth Completion
Xinjing Cheng, Peng Wang, Chenye Guan +1
cs.CVarXiv:1911.05377v22019Regularized Robust Coding for Face Recognition
Meng Yang, Lei Zhang, Jian Yang +1
cs.CVarXiv:1202.4207v22012PointCloud Saliency Maps
Tianhang Zheng, Changyou Chen, Junsong Yuan +2
cs.CVcs.AIarXiv:1812.01687v62018ECON: Explicit Clothed humans Optimized via Normal integration
Yuliang Xiu, Jinlong Yang, Xu Cao +2
cs.CVcs.AIcs.GRarXiv:2212.07422v22022Deep Flow-Guided Video Inpainting
Rui Xu, Xiaoxiao Li, Bolei Zhou +1
cs.CVarXiv:1905.02884v12019HAC: Hash-grid Assisted Context for 3D Gaussian Splatting Compression
Yihang Chen, Qianyi Wu, Weiyao Lin +2
cs.CVarXiv:2403.14530v32024Deep Gait Recognition: A Survey
Alireza Sepas-Moghaddam, Ali Etemad
cs.CVarXiv:2102.09546v22021LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation
Yixuan Ding, Jiahao Kong, Wei Huang +2
cs.CVarXiv:2608.28460v12026Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models
Simon Alexanderson, Rajmund Nagy, Jonas Beskow +1
cs.LGcs.CVcs.GRarXiv:2211.09707v22022MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild
Grégory Rogez, Cordelia Schmid
cs.CVarXiv:1607.02046v22016Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
Kai Yang, Jian Tao, Jiafei Lyu +6
cs.LGcs.AIcs.CVarXiv:2311.13231v32023SparseFusion: Distilling View-conditioned Diffusion for 3D Reconstruction
Zhizhuo Zhou, Shubham Tulsiani
cs.CVcs.GRarXiv:2212.00792v32022Interventional Few-Shot Learning
Zhongqi Yue, Hanwang Zhang, Qianru Sun +1
cs.LGcs.CVarXiv:2009.13000v22020Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution
Mostafa Dehghani, Basil Mustafa, Josip Djolonga +12
cs.CVcs.AIcs.LGarXiv:2307.06304v12023A Constructive Prediction of the Generalization Error Across Scales
Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov +1
cs.LGcs.CLcs.CVarXiv:1909.12673v22019Driver Distraction Identification with an Ensemble of Convolutional Neural Networks
Hesham M. Eraqi, Yehya Abouelnaga, Mohamed H. Saad +1
cs.CVcs.LGstat.MLarXiv:1901.09097v12019A Fully Progressive Approach to Single-Image Super-Resolution
Yifan Wang, Federico Perazzi, Brian McWilliams +3
cs.CVarXiv:1804.02900v22018Vision-based Anti-UAV Detection and Tracking
Jie Zhao, Jingshu Zhang, Dongdong Li +1
cs.CVarXiv:2205.10851v12022P+: Extended Textual Conditioning in Text-to-Image Generation
Andrey Voynov, Qinghao Chu, Daniel Cohen-Or +1
cs.CVcs.CLcs.GRarXiv:2303.09522v32023Virtual Wave Optics for Non-Line-of-Sight Imaging
Xiaochun Liu, Ibón Guillén, Marco La Manna +6
cs.CVarXiv:1810.07535v22018MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining
Xiaoyi Dong, Jianmin Bao, Yinglin Zheng +9
cs.CVarXiv:2208.12262v22022Disentangling Light Fields for Super-Resolution and Disparity Estimation
Yingqian Wang, Longguang Wang, Gaochang Wu +4
eess.IVcs.CVarXiv:2202.10603v52022Frequency Domain Model Augmentation for Adversarial Attack
Yuyang Long, Qilong Zhang, Boheng Zeng +4
cs.CVarXiv:2207.05382v12022Stacked Capsule Autoencoders
Adam R. Kosiorek, Sara Sabour, Yee Whye Teh +1
stat.MLcs.CVcs.LGarXiv:1906.06818v22019A Trilateral Weighted Sparse Coding Scheme for Real-World Image Denoising
Jun Xu, Lei Zhang, David Zhang
cs.CVarXiv:1807.04364v12018Fused DNN: A deep neural network fusion approach to fast and robust pedestrian detection
Xianzhi Du, Mostafa El-Khamy, Jungwon Lee +1
cs.CVarXiv:1610.03466v22016Robust Pre-Training by Adversarial Contrastive Learning
Ziyu Jiang, Tianlong Chen, Ting Chen +1
cs.CVarXiv:2010.13337v12020Witches' Brew: Industrial Scale Data Poisoning via Gradient Matching
Jonas Geiping, Liam Fowl, W. Ronny Huang +4
cs.CVcs.LGarXiv:2009.02276v22020SDXL-Lightning: Progressive Adversarial Diffusion Distillation
Shanchuan Lin, Anran Wang, Xiao Yang
cs.CVcs.AIcs.LGarXiv:2402.13929v32024Hybrid Spatial-Temporal Entropy Modelling for Neural Video Compression
Jiahao Li, Bin Li, Yan Lu
eess.IVcs.CVcs.MMarXiv:2207.05894v12022Agent Attention: On the Integration of Softmax and Linear Attention
Dongchen Han, Tianzhu Ye, Yizeng Han +5
cs.CVarXiv:2312.08874v32023MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying
Shaoshuai Shi, Li Jiang, Dengxin Dai +1
cs.CVarXiv:2306.17770v22023Dense Haze: A benchmark for image dehazing with dense-haze and haze-free images
Codruta O. Ancuti, Cosmin Ancuti, Mateu Sbert +1
cs.CVarXiv:1904.02904v120193D Whole Brain Segmentation using Spatially Localized Atlas Network Tiles
Yuankai Huo, Zhoubing Xu, Yunxi Xiong +7
cs.CVarXiv:1903.12152v12019Three-Dimensional Radiotherapy Dose Prediction on Head and Neck Cancer Patients with a Hierarchically Densely Connected U-net Deep Learning Architecture
Dan Nguyen, Xun Jia, David Sher +4
physics.med-phcs.CVcs.LGarXiv:1805.10397v32018