Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

16,261 to 16,320 of 18,975

  1. SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation Guidelines

    Yinda Xu, Zeyu Wang, Zuoxin Li +2

    cs.CVarXiv:1911.06188v42019
  2. The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes

    Douwe Kiela, Hamed Firooz, Aravind Mohan +4

    cs.AIcs.CLcs.CVarXiv:2005.04790v32020
  3. RMA: Rapid Motor Adaptation for Legged Robots

    Ashish Kumar, Zipeng Fu, Deepak Pathak +1

    cs.LGcs.AIcs.CVarXiv:2107.04034v12021
  4. Large Batch Training of Convolutional Networks

    Yang You, Igor Gitman, Boris Ginsburg

    cs.CVarXiv:1708.03888v32017
  5. Weakly- and Semi-Supervised Learning of a DCNN for Semantic Image Segmentation

    George Papandreou, Liang-Chieh Chen, Kevin Murphy +1

    cs.CVarXiv:1502.02734v32015
  6. A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge

    Dustin Schwenk, Apoorv Khandelwal, Christopher Clark +2

    cs.CVcs.CLarXiv:2206.01718v12022
  7. The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

    Xiaoze Liu, Ruowang Zhang, Weichen Yu +7

    cs.CLcs.CVcs.LGarXiv:2602.15382v22026
  8. MWM: Mobile World Models for Action-Conditioned Consistent Prediction

    Han Yan, Zishang Xiang, Zeyu Zhang +1

    cs.CVcs.ROarXiv:2603.07799v12026
  9. HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising

    Kai Zou, Dian Zheng, Hongbo Liu +3

    cs.CVarXiv:2603.08703v12026
  10. LIVE: Long-horizon Interactive Video World Modeling

    Junchao Huang, Ziyang Ye, Xinting Hu +5

    cs.CVarXiv:2602.03747v12026
  11. Temporal Segment Networks for Action Recognition in Videos

    Limin Wang, Yuanjun Xiong, Zhe Wang +4

    cs.CVarXiv:1705.02953v12017
  12. NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions

    Junbin Xiao, Xindi Shang, Angela Yao +1

    cs.CVcs.AIarXiv:2105.08276v22021
  13. pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis

    Eric R. Chan, Marco Monteiro, Petr Kellnhofer +2

    cs.CVcs.GRarXiv:2012.00926v22020
  14. Medical Image Segmentation Review: The success of U-Net

    Reza Azad, Ehsan Khodapanah Aghdam, Amelie Rauland +7

    eess.IVcs.CVarXiv:2211.14830v12022
  15. LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model

    Quankai Gao, Jiawei Yang, Qiangeng Xu +2

    cs.CVarXiv:2603.27449v12026
  16. Mobile-GS: Real-time Gaussian Splatting for Mobile Devices

    Xiaobiao Du, Yida Wang, Kun Zhan +1

    cs.CVarXiv:2603.11531v12026
  17. ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models

    Jooyoung Choi, Sungwon Kim, Yonghyun Jeong +2

    cs.CVarXiv:2108.02938v22021
  18. Understanding data augmentation for classification: when to warp?

    Sebastien C. Wong, Adam Gatt, Victor Stamatescu +1

    cs.CVarXiv:1609.08764v22016
  19. DeepID3: Face Recognition with Very Deep Neural Networks

    Yi Sun, Ding Liang, Xiaogang Wang +1

    cs.CVarXiv:1502.00873v12015
  20. Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

    Gen Li, Nan Duan, Yuejian Fang +3

    cs.CVarXiv:1908.06066v32019
  21. Making Convolutional Networks Shift-Invariant Again

    Richard Zhang

    cs.CVcs.LGarXiv:1904.11486v22019
  22. PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference

    Xiaofeng Mao, Shaohao Rui, Kaining Ying +4

    cs.CVcs.AIarXiv:2603.25730v12026
  23. Sparsity Invariant CNNs

    Jonas Uhrig, Nick Schneider, Lukas Schneider +3

    cs.CVarXiv:1708.06500v22017
  24. Semi-Supervised Semantic Segmentation with Cross-Consistency Training

    Yassine Ouali, Céline Hudelot, Myriam Tami

    cs.CVarXiv:2003.09005v32020
  25. MAttNet: Modular Attention Network for Referring Expression Comprehension

    Licheng Yu, Zhe Lin, Xiaohui Shen +4

    cs.CVcs.AIcs.CLarXiv:1801.08186v32018
  26. Zero-Shot Learning -- The Good, the Bad and the Ugly

    Yongqin Xian, Bernt Schiele, Zeynep Akata

    cs.CVarXiv:1703.04394v22017
  27. Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo Matching

    Xiaodong Gu, Zhiwen Fan, Zuozhuo Dai +3

    cs.CVarXiv:1912.06378v32019
  28. Deep Fragment Embeddings for Bidirectional Image Sentence Mapping

    Andrej Karpathy, Armand Joulin, Li Fei-Fei

    cs.CVcs.CLcs.LGarXiv:1406.5679v12014
  29. TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers

    Bin Yu, Shijie Lian, Xiaopeng Lin +8

    cs.ROcs.CVarXiv:2601.14133v22026
  30. MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

    Yanghao Li, Chao-Yuan Wu, Haoqi Fan +4

    cs.CVarXiv:2112.01526v22021
  31. KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3D

    Yiyi Liao, Jun Xie, Andreas Geiger

    cs.CVarXiv:2109.13410v22021
  32. BoT-SORT: Robust Associations Multi-Pedestrian Tracking

    Nir Aharon, Roy Orfaig, Ben-Zion Bobrovsky

    cs.CVarXiv:2206.14651v22022
  33. SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

    Boyuan Chen, Zhuo Xu, Sean Kirmani +6

    cs.CVcs.CLcs.LGarXiv:2401.12168v12024
  34. Repurposing Geometric Foundation Models for Multi-view Diffusion

    Wooseok Jang, Seonghu Jeon, Jisang Han +5

    cs.CVarXiv:2603.22275v12026
  35. CLIPort: What and Where Pathways for Robotic Manipulation

    Mohit Shridhar, Lucas Manuelli, Dieter Fox

    cs.ROcs.CLcs.CVarXiv:2109.12098v12021
  36. MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation

    Yejin Kim, Wilbert Pumacay, Omar Rayyan +23

    cs.ROcs.AIcs.CVarXiv:2602.11337v22026
  37. One-Shot Video Object Segmentation

    Sergi Caelles, Kevis-Kokitsi Maninis, Jordi Pont-Tuset +3

    cs.CVarXiv:1611.05198v42016
  38. Translating Videos to Natural Language Using Deep Recurrent Neural Networks

    Subhashini Venugopalan, Huijuan Xu, Jeff Donahue +3

    cs.CVcs.CLarXiv:1412.4729v32014
  39. Designing a Practical Degradation Model for Deep Blind Image Super-Resolution

    Kai Zhang, Jingyun Liang, Luc Van Gool +1

    eess.IVcs.CVarXiv:2103.14006v22021
  40. SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

    Bohao Li, Rui Wang, Guangzhi Wang +3

    cs.CLcs.CVarXiv:2307.16125v22023
  41. Temporal Action Detection with Structured Segment Networks

    Yue Zhao, Yuanjun Xiong, Limin Wang +3

    cs.CVarXiv:1704.06228v22017
  42. "Zero-Shot" Super-Resolution using Deep Internal Learning

    Assaf Shocher, Nadav Cohen, Michal Irani

    cs.CVcs.LGcs.NEarXiv:1712.06087v12017
  43. Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose

    Georgios Pavlakos, Xiaowei Zhou, Konstantinos G. Derpanis +1

    cs.CVarXiv:1611.07828v22016
  44. Leveraging Frequency Analysis for Deep Fake Image Recognition

    Joel Frank, Thorsten Eisenhofer, Lea Schönherr +3

    cs.CVeess.IVarXiv:2003.08685v32020
  45. Land-Cover Classification with High-Resolution Remote Sensing Images Using Transferable Deep Models

    Xin-Yi Tong, Gui-Song Xia, Qikai Lu +4

    cs.CVarXiv:1807.05713v32018
  46. Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models

    Meiqi Wu, Zhixin Cai, Fufangchen Zhao +13

    cs.CVarXiv:2603.22212v12026
  47. GLIGEN: Open-Set Grounded Text-to-Image Generation

    Yuheng Li, Haotian Liu, Qingyang Wu +5

    cs.CVcs.AIcs.CLarXiv:2301.07093v22023
  48. Invariant Information Clustering for Unsupervised Image Classification and Segmentation

    Xu Ji, João F. Henriques, Andrea Vedaldi

    cs.CVcs.LGarXiv:1807.06653v42018
  49. Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

    Xianjin Wu, Dingkang Liang, Tianrui Feng +5

    cs.CVcs.ROarXiv:2603.19235v32026
  50. PhyCritic: Multimodal Critic Models for Physical AI

    Tianyi Xiong, Shihao Wang, Guilin Liu +5

    cs.CVarXiv:2602.11124v12026
  51. Anomaly Detection via Reverse Distillation from One-Class Embedding

    Hanqiu Deng, Xingyu Li

    cs.CVarXiv:2201.10703v22022
  52. Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

    Keqin Chen, Zhao Zhang, Weili Zeng +3

    cs.CVarXiv:2306.15195v22023
  53. Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition

    Max Jaderberg, Karen Simonyan, Andrea Vedaldi +1

    cs.CVarXiv:1406.2227v42014
  54. Learning Continuous Image Representation with Local Implicit Image Function

    Yinbo Chen, Sifei Liu, Xiaolong Wang

    cs.CVcs.LGarXiv:2012.09161v22020
  55. Dynamic Head: Unifying Object Detection Heads with Attentions

    Xiyang Dai, Yinpeng Chen, Bin Xiao +4

    cs.CVarXiv:2106.08322v12021
  56. StrongSORT: Make DeepSORT Great Again

    Yunhao Du, Zhicheng Zhao, Yang Song +4

    cs.CVarXiv:2202.13514v22022
  57. Pros and Cons of GAN Evaluation Measures

    Ali Borji

    cs.CVarXiv:1802.03446v52018
  58. Making Reconstruction FID Predictive of Diffusion Generation FID

    Tongda Xu, Mingwei He, Shady Abu-Hussein +6

    cs.CVcs.LGarXiv:2603.05630v22026
  59. Uninformed Students: Student-Teacher Anomaly Detection with Discriminative Latent Embeddings

    Paul Bergmann, Michael Fauser, David Sattlegger +1

    cs.CVarXiv:1911.02357v22019
  60. K-Planes: Explicit Radiance Fields in Space, Time, and Appearance

    Sara Fridovich-Keil, Giacomo Meanti, Frederik Warburg +2

    cs.CVarXiv:2301.10241v22023