Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

16,081 to 16,140 of 18,839

  1. A Discriminatively Learned CNN Embedding for Person Re-identification

    Zhedong Zheng, Liang Zheng, Yi Yang

    cs.CVarXiv:1611.05666v22016
  2. Image Generation with a Sphere Encoder

    Kaiyu Yue, Menglin Jia, Ji Hou +1

    cs.CVarXiv:2602.15030v12026
  3. Hierarchical Representations for Efficient Architecture Search

    Hanxiao Liu, Karen Simonyan, Oriol Vinyals +2

    cs.LGcs.CVcs.NEarXiv:1711.00436v22017
  4. Graph R-CNN for Scene Graph Generation

    Jianwei Yang, Jiasen Lu, Stefan Lee +2

    cs.CVcs.LGarXiv:1808.00191v12018
  5. U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation

    Jun Ma, Feifei Li, Bo Wang

    eess.IVcs.CVcs.LGarXiv:2401.04722v12024
  6. Multi-Vector Index Compression in Any Modality

    Hanxiang Qin, Alexander Martin, Rohan Jha +3

    cs.IRcs.CLcs.CVarXiv:2602.21202v12026
  7. OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models

    Yufeng Zhong, Lei Chen, Xuanle Zhao +7

    cs.CVarXiv:2601.21639v22026
  8. Early Convolutions Help Transformers See Better

    Tete Xiao, Mannat Singh, Eric Mintun +3

    cs.CVarXiv:2106.14881v32021
  9. CMT: Convolutional Neural Networks Meet Vision Transformers

    Jianyuan Guo, Kai Han, Han Wu +4

    cs.CVarXiv:2107.06263v32021
  10. MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines

    Ryan Po, David Junhao Zhang, Amir Hertz +3

    cs.AIcs.CVcs.GRarXiv:2603.06679v22026
  11. UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing

    Dianyi Wang, Chaofan Ma, Feng Han +8

    cs.CVcs.AIarXiv:2602.02437v42026
  12. Medical Image Fusion: A survey of the state of the art

    A. P. James, B. V. Dasarathy

    cs.CVcs.AIphysics.med-pharXiv:1401.0166v12013
  13. How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing

    Huanyu Zhang, Xuehai Bai, Chengzu Li +9

    cs.CVarXiv:2602.01851v22026
  14. Cascaded Partial Decoder for Fast and Accurate Salient Object Detection

    Zhe Wu, Li Su, Qingming Huang

    cs.CVarXiv:1904.08739v12019
  15. OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

    Zhangquan Chen, Jiale Tao, Ruihuang Li +10

    cs.AIcs.CVarXiv:2602.05847v22026
  16. Deep Long-Tailed Learning: A Survey

    Yifan Zhang, Bingyi Kang, Bryan Hooi +2

    cs.CVarXiv:2110.04596v22021
  17. Switching Convolutional Neural Network for Crowd Counting

    Deepak Babu Sam, Shiv Surya, R. Venkatesh Babu

    cs.CVarXiv:1708.00199v22017
  18. Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition

    Yuxin Chen, Ziqi Zhang, Chunfeng Yuan +3

    cs.CVarXiv:2107.12213v22021
  19. Learning to Navigate in Complex Environments

    Piotr Mirowski, Razvan Pascanu, Fabio Viola +9

    cs.AIcs.CVcs.LGarXiv:1611.03673v32016
  20. e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings

    Haonan Chen, Sicheng Gao, Radu Timofte +2

    cs.CLcs.AIcs.CVarXiv:2601.03666v22026
  21. VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory

    Shaoan Wang, Yuanfei Luo, Xingyu Chen +6

    cs.ROcs.CVarXiv:2601.08665v12026
  22. Dynamic Neural Networks: A Survey

    Yizeng Han, Gao Huang, Shiji Song +3

    cs.CVarXiv:2102.04906v42021
  23. Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis

    Jonathon Luiten, Georgios Kopanas, Bastian Leibe +1

    cs.CVarXiv:2308.09713v12023
  24. Observation-Centric SORT: Rethinking SORT for Robust Multi-Object Tracking

    Jinkun Cao, Jiangmiao Pang, Xinshuo Weng +2

    cs.CVarXiv:2203.14360v32022
  25. Vision Transformers Need Registers

    Timothée Darcet, Maxime Oquab, Julien Mairal +1

    cs.CVarXiv:2309.16588v22023
  26. LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning

    Linquan Wu, Tianxiang Jiang, Yifei Dong +6

    cs.CVcs.AIarXiv:2601.10129v22026
  27. A Novel Focal Tversky loss function with improved Attention U-Net for lesion segmentation

    Nabila Abraham, Naimul Mefraz Khan

    cs.CVarXiv:1810.07842v12018
  28. GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

    Dhruba Ghosh, Hanna Hajishirzi, Ludwig Schmidt

    cs.CVcs.LGarXiv:2310.11513v12023
  29. PartNet: A Large-scale Benchmark for Fine-grained and Hierarchical Part-level 3D Object Understanding

    Kaichun Mo, Shilin Zhu, Angel X. Chang +4

    cs.CVarXiv:1812.02713v12018
  30. Domain Generalization by Solving Jigsaw Puzzles

    Fabio Maria Carlucci, Antonio D'Innocente, Silvia Bucci +2

    cs.CVcs.LGarXiv:1903.06864v22019
  31. Self-Refining Video Sampling

    Sangwon Jang, Taekyung Ki, Jaehyeong Jo +3

    cs.CVcs.LGarXiv:2601.18577v22026
  32. R3M: A Universal Visual Representation for Robot Manipulation

    Suraj Nair, Aravind Rajeswaran, Vikash Kumar +2

    cs.ROcs.AIcs.CVarXiv:2203.12601v32022
  33. MediX-R1: Open Ended Medical Reinforcement Learning

    Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Omair Mohamed +5

    cs.CVarXiv:2602.23363v12026
  34. ImageNet-21K Pretraining for the Masses

    Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy +1

    cs.CVcs.LGarXiv:2104.10972v42021
  35. Multi-attentional Deepfake Detection

    Hanqing Zhao, Wenbo Zhou, Dongdong Chen +3

    cs.CVarXiv:2103.02406v32021
  36. MoCha:End-to-End Video Character Replacement without Structural Guidance

    Zhengbo Xu, Jie Ma, Ziheng Wang +3

    cs.CVarXiv:2601.08587v22026
  37. SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

    Zirui Wang, Jiahui Yu, Adams Wei Yu +3

    cs.CVcs.CLcs.LGarXiv:2108.10904v32021
  38. Going Deeper in Facial Expression Recognition using Deep Neural Networks

    Ali Mollahosseini, David Chan, Mohammad H. Mahoor

    cs.NEcs.CVarXiv:1511.04110v12015
  39. Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction

    Ziyi Yang, Xinyu Gao, Wen Zhou +3

    cs.CVarXiv:2309.13101v22023
  40. Action100M: A Large-scale Video Action Dataset

    Delong Chen, Tejaswi Kasarla, Yejin Bang +6

    cs.CVarXiv:2601.10592v12026
  41. Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures

    Hengyuan Hu, Rui Peng, Yu-Wing Tai +1

    cs.NEcs.CVcs.LGarXiv:1607.03250v12016
  42. Image-Image Domain Adaptation with Preserved Self-Similarity and Domain-Dissimilarity for Person Re-identification

    Weijian Deng, Liang Zheng, Qixiang Ye +3

    cs.CVarXiv:1711.07027v32017
  43. Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

    Lijiang Li, Zuwei Long, Yunhang Shen +6

    cs.CVarXiv:2603.06577v22026
  44. UIU-Net: U-Net in U-Net for Infrared Small Object Detection

    Xin Wu, Danfeng Hong, Jocelyn Chanussot

    cs.CVarXiv:2212.00968v12022
  45. SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation Guidelines

    Yinda Xu, Zeyu Wang, Zuoxin Li +2

    cs.CVarXiv:1911.06188v42019
  46. The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes

    Douwe Kiela, Hamed Firooz, Aravind Mohan +4

    cs.AIcs.CLcs.CVarXiv:2005.04790v32020
  47. RMA: Rapid Motor Adaptation for Legged Robots

    Ashish Kumar, Zipeng Fu, Deepak Pathak +1

    cs.LGcs.AIcs.CVarXiv:2107.04034v12021
  48. Large Batch Training of Convolutional Networks

    Yang You, Igor Gitman, Boris Ginsburg

    cs.CVarXiv:1708.03888v32017
  49. Weakly- and Semi-Supervised Learning of a DCNN for Semantic Image Segmentation

    George Papandreou, Liang-Chieh Chen, Kevin Murphy +1

    cs.CVarXiv:1502.02734v32015
  50. A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge

    Dustin Schwenk, Apoorv Khandelwal, Christopher Clark +2

    cs.CVcs.CLarXiv:2206.01718v12022
  51. The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

    Xiaoze Liu, Ruowang Zhang, Weichen Yu +7

    cs.CLcs.CVcs.LGarXiv:2602.15382v22026
  52. MWM: Mobile World Models for Action-Conditioned Consistent Prediction

    Han Yan, Zishang Xiang, Zeyu Zhang +1

    cs.CVcs.ROarXiv:2603.07799v12026
  53. HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising

    Kai Zou, Dian Zheng, Hongbo Liu +3

    cs.CVarXiv:2603.08703v12026
  54. LIVE: Long-horizon Interactive Video World Modeling

    Junchao Huang, Ziyang Ye, Xinting Hu +5

    cs.CVarXiv:2602.03747v12026
  55. Temporal Segment Networks for Action Recognition in Videos

    Limin Wang, Yuanjun Xiong, Zhe Wang +4

    cs.CVarXiv:1705.02953v12017
  56. NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions

    Junbin Xiao, Xindi Shang, Angela Yao +1

    cs.CVcs.AIarXiv:2105.08276v22021
  57. pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis

    Eric R. Chan, Marco Monteiro, Petr Kellnhofer +2

    cs.CVcs.GRarXiv:2012.00926v22020
  58. Medical Image Segmentation Review: The success of U-Net

    Reza Azad, Ehsan Khodapanah Aghdam, Amelie Rauland +7

    eess.IVcs.CVarXiv:2211.14830v12022
  59. LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model

    Quankai Gao, Jiawei Yang, Qiangeng Xu +2

    cs.CVarXiv:2603.27449v12026
  60. Mobile-GS: Real-time Gaussian Splatting for Mobile Devices

    Xiaobiao Du, Yida Wang, Kun Zhan +1

    cs.CVarXiv:2603.11531v12026