Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,201 to 13,260 of 18,837

  1. Fast Dynamic Radiance Fields with Time-Aware Neural Voxels

    Jiemin Fang, Taoran Yi, Xinggang Wang +5

    cs.CVcs.GRarXiv:2205.15285v22022
  2. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

    Lin Chen, Xilin Wei, Jinsong Li +12

    cs.CVarXiv:2406.04325v12024
  3. Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos

    Serena Yeung, Olga Russakovsky, Ning Jin +3

    cs.CVarXiv:1507.05738v32015
  4. HDMapNet: An Online HD Map Construction and Evaluation Framework

    Qi Li, Yue Wang, Yilun Wang +1

    cs.CVcs.AIarXiv:2107.06307v42021
  5. Learning Not to Learn: Training Deep Neural Networks with Biased Data

    Byungju Kim, Hyunwoo Kim, Kyungsu Kim +2

    cs.CVarXiv:1812.10352v22018
  6. A scoping review of transfer learning research on medical image analysis using ImageNet

    Mohammad Amin Morid, Alireza Borjali, Guilherme Del Fiol

    eess.IVcs.CVcs.LGarXiv:2004.13175v52020
  7. Range Loss for Deep Face Recognition with Long-tail

    Xiao Zhang, Zhiyuan Fang, Yandong Wen +2

    cs.CVarXiv:1611.08976v12016
  8. Side Adapter Network for Open-Vocabulary Semantic Segmentation

    Mengde Xu, Zheng Zhang, Fangyun Wei +2

    cs.CVcs.AIarXiv:2302.12242v22023
  9. Efficient Decision-based Black-box Adversarial Attacks on Face Recognition

    Yinpeng Dong, Hang Su, Baoyuan Wu +4

    cs.CVcs.CRcs.LGarXiv:1904.04433v12019
  10. PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics

    Tianyi Xie, Zeshun Zong, Yuxing Qiu +4

    cs.GRcs.AIcs.CVarXiv:2311.12198v32023
  11. Causality matters in medical imaging

    Daniel C. Castro, Ian Walker, Ben Glocker

    eess.IVcs.AIcs.CVarXiv:1912.08142v12019
  12. Cooper: Cooperative Perception for Connected Autonomous Vehicles based on 3D Point Clouds

    Qi Chen, Sihai Tang, Qing Yang +1

    cs.CVarXiv:1905.05265v12019
  13. Learn To Pay Attention

    Saumya Jetley, Nicholas A. Lord, Namhoon Lee +1

    cs.CVcs.AIarXiv:1804.02391v22018
  14. Focal Modulation Networks

    Jianwei Yang, Chunyuan Li, Xiyang Dai +2

    cs.CVcs.AIcs.LGarXiv:2203.11926v32022
  15. A Recurrent Vision-and-Language BERT for Navigation

    Yicong Hong, Qi Wu, Yuankai Qi +2

    cs.CVarXiv:2011.13922v22020
  16. Wide-Area Image Geolocalization with Aerial Reference Imagery

    Scott Workman, Richard Souvenir, Nathan Jacobs

    cs.CVarXiv:1510.03743v12015
  17. Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

    Zhicheng Huang, Zhaoyang Zeng, Bei Liu +2

    cs.CVcs.CLcs.LGarXiv:2004.00849v22020
  18. PailitaoGR: Latent Think-with-Images for Generative Image Retrieval

    Xiaomeng Fan, Yueran Liu, Shengyu Zhou +6

    cs.CVcs.AIcs.IRarXiv:2608.26658v12026
  19. ChangeMamba: Remote Sensing Change Detection With Spatiotemporal State Space Model

    Hongruixuan Chen, Jian Song, Chengxi Han +2

    eess.IVcs.AIcs.CVarXiv:2404.03425v72024
  20. Co-Scale Conv-Attentional Image Transformers

    Weijian Xu, Yifan Xu, Tyler Chang +1

    cs.CVcs.LGcs.NEarXiv:2104.06399v22021
  21. Generative OpenMax for Multi-Class Open Set Classification

    ZongYuan Ge, Sergey Demyanov, Zetao Chen +1

    cs.CVarXiv:1707.07418v12017
  22. DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models

    Ying Fan, Olivia Watkins, Yuqing Du +7

    cs.LGcs.CVarXiv:2305.16381v32023
  23. Hierarchical Cross-Modal Talking Face Generationwith Dynamic Pixel-Wise Loss

    Lele Chen, Ross K. Maddox, Zhiyao Duan +1

    cs.CVarXiv:1905.03820v12019
  24. Beyond Sharing Weights for Deep Domain Adaptation

    Artem Rozantsev, Mathieu Salzmann, Pascal Fua

    cs.CVarXiv:1603.06432v22016
  25. Recalibrating Fully Convolutional Networks with Spatial and Channel 'Squeeze & Excitation' Blocks

    Abhijit Guha Roy, Nassir Navab, Christian Wachinger

    cs.CVarXiv:1808.08127v12018
  26. nnU-Net Revisited: A Call for Rigorous Validation in 3D Medical Image Segmentation

    Fabian Isensee, Tassilo Wald, Constantin Ulrich +4

    cs.CVarXiv:2404.09556v22024
  27. Pose-Normalized Image Generation for Person Re-identification

    Xuelin Qian, Yanwei Fu, Tao Xiang +5

    cs.CVcs.AIcs.MMarXiv:1712.02225v62017
  28. Masked Face Recognition Dataset and Application

    Zhongyuan Wang, Guangcheng Wang, Baojin Huang +11

    cs.CVarXiv:2003.09093v22020
  29. Triple Generative Adversarial Nets

    Chongxuan Li, Kun Xu, Jun Zhu +1

    cs.LGcs.CVarXiv:1703.02291v42017
  30. MixConv: Mixed Depthwise Convolutional Kernels

    Mingxing Tan, Quoc V. Le

    cs.CVcs.LGarXiv:1907.09595v32019
  31. Neural Photo Editing with Introspective Adversarial Networks

    Andrew Brock, Theodore Lim, J. M. Ritchie +1

    cs.LGcs.CVcs.NEarXiv:1609.07093v32016
  32. Image Matching across Wide Baselines: From Paper to Practice

    Yuhe Jin, Dmytro Mishkin, Anastasiia Mishchuk +4

    cs.CVarXiv:2003.01587v52020
  33. YOLO-Pose: Enhancing YOLO for Multi Person Pose Estimation Using Object Keypoint Similarity Loss

    Debapriya Maji, Soyeb Nagori, Manu Mathew +1

    cs.CVcs.AIcs.LGarXiv:2204.06806v12022
  34. Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal Mixup

    Jang-Hyun Kim, Wonho Choo, Hyun Oh Song

    cs.LGcs.AIcs.CVarXiv:2009.06962v22020
  35. UNITER: UNiversal Image-TExt Representation Learning

    Yen-Chun Chen, Linjie Li, Licheng Yu +5

    cs.CVcs.CLcs.LGarXiv:1909.11740v32019
  36. Dimensionality-Driven Learning with Noisy Labels

    Xingjun Ma, Yisen Wang, Michael E. Houle +5

    cs.CVcs.LGstat.MLarXiv:1806.02612v22018
  37. POI: Multiple Object Tracking with High Performance Detection and Appearance Feature

    Fengwei Yu, Wenbo Li, Quanquan Li +3

    cs.CVarXiv:1610.06136v12016
  38. SpinQuant: LLM quantization with learned rotations

    Zechun Liu, Changsheng Zhao, Igor Fedorov +6

    cs.LGcs.AIcs.CLarXiv:2405.16406v42024
  39. Domain Randomization and Pyramid Consistency: Simulation-to-Real Generalization without Accessing Target Domain Data

    Xiangyu Yue, Yang Zhang, Sicheng Zhao +3

    cs.CVarXiv:1909.00889v22019
  40. ESPNetv2: A Light-weight, Power Efficient, and General Purpose Convolutional Neural Network

    Sachin Mehta, Mohammad Rastegari, Linda Shapiro +1

    cs.CVarXiv:1811.11431v32018
  41. Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval

    Adam W. Harley, Alex Ufkes, Konstantinos G. Derpanis

    cs.CVcs.IRcs.LGarXiv:1502.07058v12015
  42. Message Passing Neural PDE Solvers

    Johannes Brandstetter, Daniel Worrall, Max Welling

    cs.LGcs.CVmath.NAarXiv:2202.03376v32022
  43. Procedura: Agentic 3D Modeling with Procedural Control

    Youtian Lin, Yikang Yang, Zhanpeng Hu +5

    cs.CVcs.GRarXiv:2608.26238v12026
  44. Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning

    Chenyang Wu, Fuchen Long, Binyuan Huang +4

    cs.CVcs.MMarXiv:2608.26809v12026
  45. VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance

    Katherine Crowson, Stella Biderman, Daniel Kornis +4

    cs.CVarXiv:2204.08583v22022
  46. AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection

    Qihang Zhou, Guansong Pang, Yu Tian +2

    cs.CVarXiv:2310.18961v122023
  47. GameWAM: A World Action Model for Video Games

    Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo +1

    cs.AIcs.CVcs.LGarXiv:2608.26200v12026
  48. Softmax Splatting for Video Frame Interpolation

    Simon Niklaus, Feng Liu

    cs.CVarXiv:2003.05534v12020
  49. LLaVA-Video: Video Instruction Tuning With Synthetic Data

    Yuanhan Zhang, Jinming Wu, Wei Li +4

    cs.CVcs.CLarXiv:2410.02713v32024
  50. Adversarial Reciprocal Points Learning for Open Set Recognition

    Guangyao Chen, Peixi Peng, Xiangqian Wang +1

    cs.CVarXiv:2103.00953v32021
  51. ST++: Make Self-training Work Better for Semi-supervised Semantic Segmentation

    Lihe Yang, Wei Zhuo, Lei Qi +2

    cs.CVarXiv:2106.05095v22021
  52. Multimodal Intelligence: Representation Learning, Information Fusion, and Applications

    Chao Zhang, Zichao Yang, Xiaodong He +1

    cs.AIcs.CLcs.CVarXiv:1911.03977v32019
  53. Scaling Local Self-Attention for Parameter Efficient Visual Backbones

    Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas +3

    cs.CVarXiv:2103.12731v32021
  54. Aligning Text-to-Image Models using Human Feedback

    Kimin Lee, Hao Liu, Moonkyung Ryu +6

    cs.LGcs.AIcs.CVarXiv:2302.12192v12023
  55. One Explanation Does Not Fit All: A Toolkit and Taxonomy of AI Explainability Techniques

    Vijay Arya, Rachel K. E. Bellamy, Pin-Yu Chen +17

    cs.AIcs.CVcs.HCarXiv:1909.03012v22019
  56. FloodNet: A High Resolution Aerial Imagery Dataset for Post Flood Scene Understanding

    Maryam Rahnemoonfar, Tashnim Chowdhury, Argho Sarkar +3

    cs.CVarXiv:2012.02951v12020
  57. UniDepth: Universal Monocular Metric Depth Estimation

    Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis +4

    cs.CVarXiv:2403.18913v12024
  58. Gauge Equivariant Convolutional Networks and the Icosahedral CNN

    Taco S. Cohen, Maurice Weiler, Berkay Kicanaoglu +1

    cs.LGcs.CVcs.NEarXiv:1902.04615v32019
  59. Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation Learning

    Zhenda Xie, Yutong Lin, Zheng Zhang +3

    cs.CVcs.LGarXiv:2011.10043v22020
  60. Stochastic Adversarial Video Prediction

    Alex X. Lee, Richard Zhang, Frederik Ebert +3

    cs.CVcs.AIcs.LGarXiv:1804.01523v12018