Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

11,101 to 11,160 of 18,839

  1. ForensicTransfer: Weakly-supervised Domain Adaptation for Forgery Detection

    Davide Cozzolino, Justus Thies, Andreas Rössler +3

    cs.CVarXiv:1812.02510v22018
  2. LDSO: Direct Sparse Odometry with Loop Closure

    Xiang Gao, Rui Wang, Nikolaus Demmel +1

    cs.CVarXiv:1808.01111v12018
  3. Learning Dynamic Memory Networks for Object Tracking

    Tianyu Yang, Antoni B. Chan

    cs.CVarXiv:1803.07268v22018
  4. VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning

    Bo Jiang, Shaoyu Chen, Hao Gao +4

    cs.CVcs.ROarXiv:2402.13243v22024
  5. ManiGAN: Text-Guided Image Manipulation

    Bowen Li, Xiaojuan Qi, Thomas Lukasiewicz +1

    cs.CVcs.CLcs.LGarXiv:1912.06203v22019
  6. Disentangling Label Distribution for Long-tailed Visual Recognition

    Youngkyu Hong, Seungju Han, Kwanghee Choi +3

    cs.CVcs.LGarXiv:2012.00321v22020
  7. Towards Faster and Stabilized GAN Training for High-fidelity Few-shot Image Synthesis

    Bingchen Liu, Yizhe Zhu, Kunpeng Song +1

    cs.CVcs.AIarXiv:2101.04775v12021
  8. DiffusionNet: Discretization Agnostic Learning on Surfaces

    Nicholas Sharp, Souhaib Attaiki, Keenan Crane +1

    cs.CVcs.CGcs.LGarXiv:2012.00888v32020
  9. Robust and Precise Vehicle Localization based on Multi-sensor Fusion in Diverse City Scenes

    Guowei Wan, Xiaolong Yang, Renlan Cai +3

    cs.CVcs.ROarXiv:1711.05805v22017
  10. Self-Monitoring Navigation Agent via Auxiliary Progress Estimation

    Chih-Yao Ma, Jiasen Lu, Zuxuan Wu +4

    cs.AIcs.CLcs.CVarXiv:1901.03035v12019
  11. Deep Dense Multi-scale Network for Snow Removal Using Semantic and Geometric Priors

    Kaihao Zhang, Rongqing Li, Yanjiang Yu +3

    cs.CVarXiv:2103.11298v12021
  12. Perpetual Humanoid Control for Real-time Simulated Avatars

    Zhengyi Luo, Jinkun Cao, Alexander Winkler +2

    cs.CVcs.GRcs.ROarXiv:2305.06456v32023
  13. Unsupervised Attention-guided Image to Image Translation

    Youssef A. Mejjati, Christian Richardt, James Tompkin +2

    cs.CVcs.AIarXiv:1806.02311v32018
  14. Omnivore: A Single Model for Many Visual Modalities

    Rohit Girdhar, Mannat Singh, Nikhila Ravi +3

    cs.CVcs.AIcs.IRarXiv:2201.08377v22022
  15. Weakly Supervised Object Localization and Detection: A Survey

    Dingwen Zhang, Junwei Han, Gong Cheng +1

    cs.CVarXiv:2104.07918v12021
  16. Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation

    Dan Xu, Wei Wang, Hao Tang +3

    cs.CVarXiv:1803.11029v12018
  17. Learning to Assemble Neural Module Tree Networks for Visual Grounding

    Daqing Liu, Hanwang Zhang, Feng Wu +1

    cs.CVarXiv:1812.03299v32018
  18. Fully Connected Deep Structured Networks

    Alexander G. Schwing, Raquel Urtasun

    cs.CVcs.LGarXiv:1503.02351v12015
  19. Learning Robust Features using Deep Learning for Automatic Seizure Detection

    Pierre Thodoroff, Joelle Pineau, Andrew Lim

    cs.LGcs.CVarXiv:1608.00220v12016
  20. LeViT-UNet: Make Faster Encoders with Transformer for Medical Image Segmentation

    Guoping Xu, Xingrong Wu, Xuan Zhang +1

    cs.CVarXiv:2107.08623v12021
  21. Optical Flow Guided Feature: A Fast and Robust Motion Representation for Video Action Recognition

    Shuyang Sun, Zhanghui Kuang, Wanli Ouyang +2

    cs.CVarXiv:1711.11152v22017
  22. Deep Self-Learning From Noisy Labels

    Jiangfan Han, Ping Luo, Xiaogang Wang

    cs.CVcs.LGarXiv:1908.02160v22019
  23. VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction

    Jiaqi Lin, Zhihao Li, Xiao Tang +8

    cs.CVarXiv:2402.17427v12024
  24. What's "up" with vision-language models? Investigating their struggle with spatial reasoning

    Amita Kamath, Jack Hessel, Kai-Wei Chang

    cs.CLcs.CVcs.LGarXiv:2310.19785v12023
  25. Feature-metric Registration: A Fast Semi-supervised Approach for Robust Point Cloud Registration without Correspondences

    Xiaoshui Huang, Guofeng Mei, Jian Zhang

    cs.CVarXiv:2005.01014v12020
  26. Domain Adaptation for Semantic Segmentation with Maximum Squares Loss

    Minghao Chen, Hongyang Xue, Deng Cai

    cs.CVarXiv:1909.13589v12019
  27. Self-Supervised GANs via Auxiliary Rotation Loss

    Ting Chen, Xiaohua Zhai, Marvin Ritter +2

    cs.LGcs.CVstat.MLarXiv:1811.11212v22018
  28. Deep Matching Prior Network: Toward Tighter Multi-oriented Text Detection

    Yuliang Liu, Lianwen Jin

    cs.CVarXiv:1703.01425v12017
  29. Photographic Text-to-Image Synthesis with a Hierarchically-nested Adversarial Network

    Zizhao Zhang, Yuanpu Xie, Lin Yang

    cs.CVarXiv:1802.09178v22018
  30. GaussianShader: 3D Gaussian Splatting with Shading Functions for Reflective Surfaces

    Yingwenqi Jiang, Jiadong Tu, Yuan Liu +4

    cs.CVarXiv:2311.17977v12023
  31. PLOP: Learning without Forgetting for Continual Semantic Segmentation

    Arthur Douillard, Yifu Chen, Arnaud Dapogny +1

    cs.CVarXiv:2011.11390v32020
  32. DeFRCN: Decoupled Faster R-CNN for Few-Shot Object Detection

    Limeng Qiao, Yuxuan Zhao, Zhiyuan Li +3

    cs.CVarXiv:2108.09017v12021
  33. CAT: Cross Attention in Vision Transformer

    Hezheng Lin, Xing Cheng, Xiangyu Wu +5

    cs.CVcs.AIarXiv:2106.05786v12021
  34. GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models

    Taoran Yi, Jiemin Fang, Junjie Wang +6

    cs.CVcs.GRarXiv:2310.08529v32023
  35. Deep Graph Clustering via Dual Correlation Reduction

    Yue Liu, Wenxuan Tu, Sihang Zhou +4

    cs.LGcs.AIcs.CVarXiv:2112.14772v12021
  36. SIXray : A Large-scale Security Inspection X-ray Benchmark for Prohibited Item Discovery in Overlapping Images

    Caijing Miao, Lingxi Xie, Fang Wan +4

    cs.CVarXiv:1901.00303v12019
  37. Generating Images with Sparse Representations

    Charlie Nash, Jacob Menick, Sander Dieleman +1

    cs.CVstat.MLarXiv:2103.03841v12021
  38. Deep Residual Learning for Accelerated MRI using Magnitude and Phase Networks

    Dongwook Lee, Jaejun Yoo, Sungho Tak +1

    cs.CVcs.AIcs.LGarXiv:1804.00432v12018
  39. Model-free Deep Reinforcement Learning for Urban Autonomous Driving

    Jianyu Chen, Bodi Yuan, Masayoshi Tomizuka

    cs.LGcs.AIcs.CVarXiv:1904.09503v22019
  40. Fusion of Multispectral Data Through Illumination-aware Deep Neural Networks for Pedestrian Detection

    Dayan Guan, Yanpeng Cao, Jun Liang +2

    cs.CVarXiv:1802.09972v12018
  41. Social Ways: Learning Multi-Modal Distributions of Pedestrian Trajectories with GANs

    Javad Amirian, Jean-Bernard Hayet, Julien Pettre

    cs.CVarXiv:1904.09507v22019
  42. CAT2000: A Large Scale Fixation Dataset for Boosting Saliency Research

    Ali Borji, Laurent Itti

    cs.CVarXiv:1505.03581v12015
  43. CoEdge: Cooperative DNN Inference with Adaptive Workload Partitioning over Heterogeneous Edge Devices

    Liekang Zeng, Xu Chen, Zhi Zhou +2

    cs.NIcs.CVcs.DCarXiv:2012.03257v12020
  44. Seq-NMS for Video Object Detection

    Wei Han, Pooya Khorrami, Tom Le Paine +6

    cs.CVarXiv:1602.08465v32016
  45. MonoViT: Self-Supervised Monocular Depth Estimation with a Vision Transformer

    Chaoqiang Zhao, Youmin Zhang, Matteo Poggi +6

    cs.CVarXiv:2208.03543v12022
  46. Learning Salient Boundary Feature for Anchor-free Temporal Action Localization

    Chuming Lin, Chengming Xu, Donghao Luo +6

    cs.CVcs.AIarXiv:2103.13137v12021
  47. Few-shot Image Generation via Cross-domain Correspondence

    Utkarsh Ojha, Yijun Li, Jingwan Lu +4

    cs.CVcs.GRcs.LGarXiv:2104.06820v12021
  48. Lip Movements Generation at a Glance

    Lele Chen, Zhiheng Li, Ross K. Maddox +2

    cs.CVarXiv:1803.10404v32018
  49. Deep Generalized Unfolding Networks for Image Restoration

    Chong Mou, Qian Wang, Jian Zhang

    cs.CVeess.IVarXiv:2204.13348v12022
  50. Style Aggregated Network for Facial Landmark Detection

    Xuanyi Dong, Yan Yan, Wanli Ouyang +1

    cs.CVarXiv:1803.04108v42018
  51. Facial Expression Recognition with Visual Transformers and Attentional Selective Fusion

    Fuyan Ma, Bin Sun, Shutao Li

    cs.CVarXiv:2103.16854v32021
  52. Slow Feature Analysis for Human Action Recognition

    Zhang Zhang, Dacheng Tao

    cs.CVarXiv:1907.06670v12019
  53. Recurrent Human Pose Estimation

    Vasileios Belagiannis, Andrew Zisserman

    cs.CVcs.NEarXiv:1605.02914v32016
  54. Infrared and Visible Image Fusion with ResNet and zero-phase component analysis

    Hui Li, Xiao-Jun Wu, Tariq S. Durrani

    cs.CVarXiv:1806.07119v72018
  55. Shunted Self-Attention via Multi-Scale Token Aggregation

    Sucheng Ren, Daquan Zhou, Shengfeng He +2

    cs.CVarXiv:2111.15193v22021
  56. Open-Vocabulary DETR with Conditional Matching

    Yuhang Zang, Wei Li, Kaiyang Zhou +2

    cs.CVcs.AIarXiv:2203.11876v22022
  57. Learning Regularity in Skeleton Trajectories for Anomaly Detection in Videos

    Romero Morais, Vuong Le, Truyen Tran +3

    cs.CVarXiv:1903.03295v22019
  58. Supervised Multimodal Bitransformers for Classifying Images and Text

    Douwe Kiela, Suvrat Bhooshan, Hamed Firooz +2

    cs.CLcs.CVcs.LGarXiv:1909.02950v22019
  59. Visual Camera Re-Localization from RGB and RGB-D Images Using DSAC

    Eric Brachmann, Carsten Rother

    cs.CVcs.LGarXiv:2002.12324v42020
  60. Patch-based Progressive 3D Point Set Upsampling

    Wang Yifan, Shihao Wu, Hui Huang +2

    cs.CVcs.GRcs.LGarXiv:1811.11286v32018