Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

10,141 to 10,200 of 18,867

  1. Fast Alternating Linearization Methods for Minimizing the Sum of Two Convex Functions

    Donald Goldfarb, Shiqian Ma, Katya Scheinberg

    math.OCcs.CVmath.NAarXiv:0912.4571v22009
  2. Beyond Temporal Pooling: Recurrence and Temporal Convolutions for Gesture Recognition in Video

    Lionel Pigou, Aäron van den Oord, Sander Dieleman +2

    cs.CVcs.AIcs.LGarXiv:1506.01911v32015
  3. Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

    Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang +2

    cs.CVarXiv:2608.28192v12026
  4. Monte Carlo Convolution for Learning on Non-Uniformly Sampled Point Clouds

    Pedro Hermosilla, Tobias Ritschel, Pere-Pau Vázquez +2

    cs.CVarXiv:1806.01759v22018
  5. Density Map Guided Object Detection in Aerial Images

    Changlin Li, Taojiannan Yang, Sijie Zhu +2

    cs.CVarXiv:2004.05520v12020
  6. Cross-domain Detection via Graph-induced Prototype Alignment

    Minghao Xu, Hang Wang, Bingbing Ni +2

    cs.CVarXiv:2003.12849v12020
  7. SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation

    Wenchao Sun, Xuewu Lin, Yining Shi +3

    cs.CVarXiv:2405.19620v22024
  8. Person Re-identification by Contour Sketch under Moderate Clothing Change

    Qize Yang, Ancong Wu, Wei-Shi Zheng

    cs.CVarXiv:2002.02295v12020
  9. Cut and Learn for Unsupervised Object Detection and Instance Segmentation

    Xudong Wang, Rohit Girdhar, Stella X. Yu +1

    cs.CVcs.AIcs.LGarXiv:2301.11320v12023
  10. 3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans

    Antoni Rosinol, Arjun Gupta, Marcus Abate +2

    cs.ROcs.AIcs.CVarXiv:2002.06289v22020
  11. Uncertainty-aware multi-view co-training for semi-supervised medical image segmentation and domain adaptation

    Yingda Xia, Dong Yang, Zhiding Yu +7

    cs.CVarXiv:2006.16806v12020
  12. GP-GAN: Towards Realistic High-Resolution Image Blending

    Huikai Wu, Shuai Zheng, Junge Zhang +1

    cs.CVarXiv:1703.07195v32017
  13. A Survey of Vision-Language Pre-Trained Models

    Yifan Du, Zikang Liu, Junyi Li +1

    cs.CVcs.CLcs.LGarXiv:2202.10936v22022
  14. Automatic 3D liver location and segmentation via convolutional neural networks and graph cut

    Fang Lu, Fa Wu, Peijun Hu +2

    cs.CVarXiv:1605.03012v12016
  15. GS-IR: 3D Gaussian Splatting for Inverse Rendering

    Zhihao Liang, Qi Zhang, Ying Feng +2

    cs.CVarXiv:2311.16473v32023
  16. Improving Vision-and-Language Navigation with Image-Text Pairs from the Web

    Arjun Majumdar, Ayush Shrivastava, Stefan Lee +3

    cs.CVcs.AIcs.CLarXiv:2004.14973v22020
  17. Beyond the Pixel-Wise Loss for Topology-Aware Delineation

    Agata Mosinska, Pablo Marquez-Neila, Mateusz Kozinski +1

    cs.CVarXiv:1712.02190v12017
  18. Robust and Generalizable Visual Representation Learning via Random Convolutions

    Zhenlin Xu, Deyi Liu, Junlin Yang +2

    cs.CVcs.LGarXiv:2007.13003v32020
  19. Learning a Deep ConvNet for Multi-label Classification with Partial Labels

    Thibaut Durand, Nazanin Mehrasa, Greg Mori

    cs.CVarXiv:1902.09720v12019
  20. LAPAR: Linearly-Assembled Pixel-Adaptive Regression Network for Single Image Super-Resolution and Beyond

    Wenbo Li, Kun Zhou, Lu Qi +3

    cs.CVarXiv:2105.10422v12021
  21. Exploring Spatial Context for 3D Semantic Segmentation of Point Clouds

    Francis Engelmann, Theodora Kontogianni, Alexander Hermans +1

    cs.CVarXiv:1802.01500v22018
  22. LMSCNet: Lightweight Multiscale 3D Semantic Completion

    Luis Roldão, Raoul de Charette, Anne Verroust-Blondet

    cs.CVarXiv:2008.10559v22020
  23. Scene as Occupancy

    Chonghao Sima, Wenwen Tong, Tai Wang +8

    cs.CVcs.ROarXiv:2306.02851v32023
  24. ReNet: A Recurrent Neural Network Based Alternative to Convolutional Networks

    Francesco Visin, Kyle Kastner, Kyunghyun Cho +3

    cs.CVarXiv:1505.00393v32015
  25. Guided Motion Diffusion for Controllable Human Motion Synthesis

    Korrawe Karunratanakul, Konpat Preechakul, Supasorn Suwajanakorn +1

    cs.CVarXiv:2305.12577v32023
  26. CSPN++: Learning Context and Resource Aware Convolutional Spatial Propagation Networks for Depth Completion

    Xinjing Cheng, Peng Wang, Chenye Guan +1

    cs.CVarXiv:1911.05377v22019
  27. Regularized Robust Coding for Face Recognition

    Meng Yang, Lei Zhang, Jian Yang +1

    cs.CVarXiv:1202.4207v22012
  28. PointCloud Saliency Maps

    Tianhang Zheng, Changyou Chen, Junsong Yuan +2

    cs.CVcs.AIarXiv:1812.01687v62018
  29. ECON: Explicit Clothed humans Optimized via Normal integration

    Yuliang Xiu, Jinlong Yang, Xu Cao +2

    cs.CVcs.AIcs.GRarXiv:2212.07422v22022
  30. Deep Flow-Guided Video Inpainting

    Rui Xu, Xiaoxiao Li, Bolei Zhou +1

    cs.CVarXiv:1905.02884v12019
  31. HAC: Hash-grid Assisted Context for 3D Gaussian Splatting Compression

    Yihang Chen, Qianyi Wu, Weiyao Lin +2

    cs.CVarXiv:2403.14530v32024
  32. Deep Gait Recognition: A Survey

    Alireza Sepas-Moghaddam, Ali Etemad

    cs.CVarXiv:2102.09546v22021
  33. LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation

    Yixuan Ding, Jiahao Kong, Wei Huang +2

    cs.CVarXiv:2608.28460v12026
  34. Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models

    Simon Alexanderson, Rajmund Nagy, Jonas Beskow +1

    cs.LGcs.CVcs.GRarXiv:2211.09707v22022
  35. MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild

    Grégory Rogez, Cordelia Schmid

    cs.CVarXiv:1607.02046v22016
  36. Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model

    Kai Yang, Jian Tao, Jiafei Lyu +6

    cs.LGcs.AIcs.CVarXiv:2311.13231v32023
  37. SparseFusion: Distilling View-conditioned Diffusion for 3D Reconstruction

    Zhizhuo Zhou, Shubham Tulsiani

    cs.CVcs.GRarXiv:2212.00792v32022
  38. Interventional Few-Shot Learning

    Zhongqi Yue, Hanwang Zhang, Qianru Sun +1

    cs.LGcs.CVarXiv:2009.13000v22020
  39. Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution

    Mostafa Dehghani, Basil Mustafa, Josip Djolonga +12

    cs.CVcs.AIcs.LGarXiv:2307.06304v12023
  40. A Constructive Prediction of the Generalization Error Across Scales

    Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov +1

    cs.LGcs.CLcs.CVarXiv:1909.12673v22019
  41. Driver Distraction Identification with an Ensemble of Convolutional Neural Networks

    Hesham M. Eraqi, Yehya Abouelnaga, Mohamed H. Saad +1

    cs.CVcs.LGstat.MLarXiv:1901.09097v12019
  42. A Fully Progressive Approach to Single-Image Super-Resolution

    Yifan Wang, Federico Perazzi, Brian McWilliams +3

    cs.CVarXiv:1804.02900v22018
  43. Vision-based Anti-UAV Detection and Tracking

    Jie Zhao, Jingshu Zhang, Dongdong Li +1

    cs.CVarXiv:2205.10851v12022
  44. P+: Extended Textual Conditioning in Text-to-Image Generation

    Andrey Voynov, Qinghao Chu, Daniel Cohen-Or +1

    cs.CVcs.CLcs.GRarXiv:2303.09522v32023
  45. Virtual Wave Optics for Non-Line-of-Sight Imaging

    Xiaochun Liu, Ibón Guillén, Marco La Manna +6

    cs.CVarXiv:1810.07535v22018
  46. MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining

    Xiaoyi Dong, Jianmin Bao, Yinglin Zheng +9

    cs.CVarXiv:2208.12262v22022
  47. Disentangling Light Fields for Super-Resolution and Disparity Estimation

    Yingqian Wang, Longguang Wang, Gaochang Wu +4

    eess.IVcs.CVarXiv:2202.10603v52022
  48. Frequency Domain Model Augmentation for Adversarial Attack

    Yuyang Long, Qilong Zhang, Boheng Zeng +4

    cs.CVarXiv:2207.05382v12022
  49. Stacked Capsule Autoencoders

    Adam R. Kosiorek, Sara Sabour, Yee Whye Teh +1

    stat.MLcs.CVcs.LGarXiv:1906.06818v22019
  50. A Trilateral Weighted Sparse Coding Scheme for Real-World Image Denoising

    Jun Xu, Lei Zhang, David Zhang

    cs.CVarXiv:1807.04364v12018
  51. Fused DNN: A deep neural network fusion approach to fast and robust pedestrian detection

    Xianzhi Du, Mostafa El-Khamy, Jungwon Lee +1

    cs.CVarXiv:1610.03466v22016
  52. Robust Pre-Training by Adversarial Contrastive Learning

    Ziyu Jiang, Tianlong Chen, Ting Chen +1

    cs.CVarXiv:2010.13337v12020
  53. Witches' Brew: Industrial Scale Data Poisoning via Gradient Matching

    Jonas Geiping, Liam Fowl, W. Ronny Huang +4

    cs.CVcs.LGarXiv:2009.02276v22020
  54. SDXL-Lightning: Progressive Adversarial Diffusion Distillation

    Shanchuan Lin, Anran Wang, Xiao Yang

    cs.CVcs.AIcs.LGarXiv:2402.13929v32024
  55. Hybrid Spatial-Temporal Entropy Modelling for Neural Video Compression

    Jiahao Li, Bin Li, Yan Lu

    eess.IVcs.CVcs.MMarXiv:2207.05894v12022
  56. Agent Attention: On the Integration of Softmax and Linear Attention

    Dongchen Han, Tianzhu Ye, Yizeng Han +5

    cs.CVarXiv:2312.08874v32023
  57. MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying

    Shaoshuai Shi, Li Jiang, Dengxin Dai +1

    cs.CVarXiv:2306.17770v22023
  58. Dense Haze: A benchmark for image dehazing with dense-haze and haze-free images

    Codruta O. Ancuti, Cosmin Ancuti, Mateu Sbert +1

    cs.CVarXiv:1904.02904v12019
  59. 3D Whole Brain Segmentation using Spatially Localized Atlas Network Tiles

    Yuankai Huo, Zhoubing Xu, Yunxi Xiong +7

    cs.CVarXiv:1903.12152v12019
  60. Three-Dimensional Radiotherapy Dose Prediction on Head and Neck Cancer Patients with a Hierarchically Densely Connected U-net Deep Learning Architecture

    Dan Nguyen, Xun Jia, David Sher +4

    physics.med-phcs.CVcs.LGarXiv:1805.10397v32018