Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,601 to 3,660 of 18,796

  1. Fast Resampling of 3D Point Clouds via Graphs

    Siheng Chen, Dong Tian, Chen Feng +2

    cs.CVarXiv:1702.06397v12017
  2. Bottom Up Top Down Detection Transformers for Language Grounding in Images and Point Clouds

    Ayush Jain, Nikolaos Gkanatsios, Ishita Mediratta +1

    cs.CVcs.CLarXiv:2112.08879v52021
  3. Count-ception: Counting by Fully Convolutional Redundant Counting

    Joseph Paul Cohen, Genevieve Boucher, Craig A. Glastonbury +2

    cs.CVcs.LGstat.MLarXiv:1703.08710v22017
  4. Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing

    Hao Yang, Zhiyu Tan, Jia Gong +7

    cs.CVarXiv:2602.08820v22026
  5. Spatial Pyramid Based Graph Reasoning for Semantic Segmentation

    Xia Li, Yibo Yang, Qijie Zhao +3

    cs.CVarXiv:2003.10211v12020
  6. L2G: A Simple Local-to-Global Knowledge Transfer Framework for Weakly Supervised Semantic Segmentation

    Peng-Tao Jiang, Yuqi Yang, Qibin Hou +1

    cs.CVarXiv:2204.03206v12022
  7. Physics-based Noise Modeling for Extreme Low-light Photography

    Kaixuan Wei, Ying Fu, Yinqiang Zheng +1

    eess.IVcs.CVarXiv:2108.02158v12021
  8. MIRROR: Manifold Ideal Reference ReconstructOR for Generalizable AI-Generated Image Detection

    Ruiqi Liu, Manni Cui, Ziheng Qin +12

    cs.CVcs.CRarXiv:2602.02222v12026
  9. Provable Dynamic Fusion for Low-Quality Multimodal Data

    Qingyang Zhang, Haitao Wu, Changqing Zhang +4

    cs.LGcs.CVarXiv:2306.02050v22023
  10. MV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D Generation

    Baicheng Li, Dong Wu, Jun Li +4

    cs.CVarXiv:2603.11633v22026
  11. Clustering Millions of Faces by Identity

    Charles Otto, Dayong Wang, Anil K. Jain

    cs.CVarXiv:1604.00989v12016
  12. Zero-Shot Grounding of Objects from Natural Language Queries

    Arka Sadhu, Kan Chen, Ram Nevatia

    cs.CVcs.CLarXiv:1908.07129v12019
  13. Fast Low-rank Shared Dictionary Learning for Image Classification

    Tiep Vu, Vishal Monga

    cs.CVcs.AIarXiv:1610.08606v32016
  14. Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models

    Hao Tang, Yu Liu, Shuanglin Yan +3

    cs.CVcs.MMarXiv:2601.08476v22026
  15. Contour-Guided Query-Based Feature Fusion for Boundary-Aware and Generalizable Cardiac Ultrasound Segmentation

    Zahid Ullah, Sieun Choi, Jihie Kim

    cs.CVarXiv:2603.28110v12026
  16. ReRoPE: Repurposing RoPE for Relative Camera Control

    Chunyang Li, Yuanbo Yang, Jiahao Shao +3

    cs.CVarXiv:2602.08068v12026
  17. Learning Fast, Learning Slow: A General Continual Learning Method based on Complementary Learning System

    Elahe Arani, Fahad Sarfraz, Bahram Zonooz

    cs.LGcs.AIcs.CVarXiv:2201.12604v22022
  18. DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving

    Shuyao Shang, Bing Zhan, Yunfei Yan +9

    cs.CVcs.ROarXiv:2603.11041v22026
  19. Diverse Weight Averaging for Out-of-Distribution Generalization

    Alexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier +3

    cs.CVcs.AIcs.LGarXiv:2205.09739v22022
  20. X-View: Graph-Based Semantic Multi-View Localization

    Abel Gawel, Carlo Del Don, Roland Siegwart +2

    cs.ROcs.CVarXiv:1709.09905v32017
  21. TC-SSA: Token Compression via Semantic Slot Aggregation for Gigapixel Pathology Reasoning

    Zhuo Chen, Shawn Young, Lijian Xu

    cs.CVcs.AIarXiv:2603.01143v12026
  22. Fully Convolutional Crowd Counting On Highly Congested Scenes

    Mark Marsden, Kevin McGuinness, Suzanne Little +1

    cs.CVarXiv:1612.00220v22016
  23. Lifelong Machine Learning with Deep Streaming Linear Discriminant Analysis

    Tyler L. Hayes, Christopher Kanan

    cs.LGcs.CVstat.MLarXiv:1909.01520v32019
  24. Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models

    Zhiqiu Lin, Samuel Yu, Zhiyi Kuang +2

    cs.CVcs.AIcs.LGarXiv:2301.06267v52023
  25. When NAS Meets Robustness: In Search of Robust Architectures against Adversarial Attacks

    Minghao Guo, Yuzhe Yang, Rui Xu +2

    cs.LGcs.CRcs.CVarXiv:1911.10695v32019
  26. Trans-SVNet: Accurate Phase Recognition from Surgical Videos via Hybrid Embedding Aggregation Transformer

    Xiaojie Gao, Yueming Jin, Yonghao Long +2

    cs.CVcs.AIarXiv:2103.09712v22021
  27. TransCenter: Transformers with Dense Representations for Multiple-Object Tracking

    Yihong Xu, Yutong Ban, Guillaume Delorme +3

    cs.CVarXiv:2103.15145v42021
  28. HF-NeuS: Improved Surface Reconstruction Using High-Frequency Details

    Yiqun Wang, Ivan Skorokhodov, Peter Wonka

    cs.CVcs.GRarXiv:2206.07850v22022
  29. Visual Question Generation as Dual Task of Visual Question Answering

    Yikang Li, Nan Duan, Bolei Zhou +3

    cs.CVarXiv:1709.07192v12017
  30. SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation

    Jiwen Zhang, Zejun Li, Siyuan Wang +3

    cs.CVcs.AIcs.ROarXiv:2601.06806v12026
  31. Context-Aware Trajectory Prediction

    Federico Bartoli, Giuseppe Lisanti, Lamberto Ballan +1

    cs.CVarXiv:1705.02503v12017
  32. TF-Blender: Temporal Feature Blender for Video Object Detection

    Yiming Cui, Liqi Yan, Zhiwen Cao +1

    cs.CVarXiv:2108.05821v12021
  33. ARM: Advantage Reward Modeling for Long-Horizon Manipulation

    Yiming Mao, Zixi Yu, Weixin Mao +5

    cs.ROcs.AIcs.CVarXiv:2604.03037v22026
  34. NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation

    Huichao Zhang, Liao Qu, Yiheng Liu +33

    cs.CVcs.AIarXiv:2601.02204v12026
  35. 3D Human Pose Estimation in RGBD Images for Robotic Task Learning

    Christian Zimmermann, Tim Welschehold, Christian Dornhege +2

    cs.CVcs.ROarXiv:1803.02622v22018
  36. AutoFigure-Edit: Generating Editable Scientific Illustration

    Zhen Lin, Qiujie Xie, Minjun Zhu +10

    cs.CVcs.AIarXiv:2603.06674v12026
  37. Synthetic data generation for end-to-end thermal infrared tracking

    Lichao Zhang, Abel Gonzalez-Garcia, Joost van de Weijer +2

    cs.CVarXiv:1806.01013v22018
  38. Landslide4Sense: Reference Benchmark Data and Deep Learning Models for Landslide Detection

    Omid Ghorbanzadeh, Yonghao Xu, Pedram Ghamisi +2

    cs.CVeess.IVarXiv:2206.00515v32022
  39. SynQ: Accurate Zero-shot Quantization by Synthesis-aware Fine-tuning

    Minjun Kim, Jongjin Kim, U Kang

    cs.CVarXiv:2603.18423v12026
  40. Prostate Cancer Diagnosis using Deep Learning with 3D Multiparametric MRI

    Saifeng Liu, Huaixiu Zheng, Yesu Feng +1

    cs.CVstat.MLarXiv:1703.04078v12017
  41. Hierarchical Dense Correlation Distillation for Few-Shot Segmentation

    Bohao Peng, Zhuotao Tian, Xiaoyang Wu +4

    cs.CVcs.AIarXiv:2303.14652v12023
  42. AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots

    Likui Zhang, Tao Tang, Zhihao Zhan +9

    cs.ROcs.AIcs.CVarXiv:2603.07648v12026
  43. Self-Supervised Learning of Audio-Visual Objects from Video

    Triantafyllos Afouras, Andrew Owens, Joon Son Chung +1

    cs.CVcs.SDeess.ASarXiv:2008.04237v12020
  44. Online Video Deblurring via Dynamic Temporal Blending Network

    Tae Hyun Kim, Kyoung Mu Lee, Bernhard Schölkopf +1

    cs.CVarXiv:1704.03285v12017
  45. Full-Duplex Strategy for Video Object Segmentation

    Ge-Peng Ji, Deng-Ping Fan, Keren Fu +3

    cs.CVarXiv:2108.03151v32021
  46. LoMa: Local Feature Matching Revisited

    David Nordström, Johan Edstedt, Georg Bökman +6

    cs.CVarXiv:2604.04931v22026
  47. Deep Gaussian Scale Mixture Prior for Spectral Compressive Imaging

    Tao Huang, Weisheng Dong, Xin Yuan +2

    eess.IVcs.CVarXiv:2103.07152v22021
  48. Point Cloud Upsampling via Disentangled Refinement

    Ruihui Li, Xianzhi Li, Pheng-Ann Heng +1

    cs.CVarXiv:2106.04779v12021
  49. TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention

    Chuancheng Shi, Shangze Li, Wenjun Lu +5

    cs.CVcs.AIcs.CYarXiv:2601.21900v22026
  50. Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting

    Hao Feng, Wei Shi, Ke Zhang +9

    cs.CVarXiv:2602.05384v12026
  51. Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

    Lili Yu, Bowen Shi, Ramakanth Pasunuru +24

    cs.LGcs.CLcs.CVarXiv:2309.02591v12023
  52. PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction

    Peng Wang, Hao Tan, Sai Bi +6

    cs.CVarXiv:2311.12024v22023
  53. Deep DIC: Deep Learning-Based Digital Image Correlation for End-to-End Displacement and Strain Measurement

    Ru Yang, Yang Li, Danielle Zeng +1

    eess.IVcond-mat.mtrl-scics.CVarXiv:2110.13720v22021
  54. Contrastive and Non-Contrastive Self-Supervised Learning Recover Global and Local Spectral Embedding Methods

    Randall Balestriero, Yann LeCun

    cs.LGcs.AIcs.CVarXiv:2205.11508v32022
  55. RenderGAN: Generating Realistic Labeled Data

    Leon Sixt, Benjamin Wild, Tim Landgraf

    cs.NEcs.CVarXiv:1611.01331v52016
  56. Visually-Guided Policy Optimization for Multimodal Reasoning

    Zengbin Wang, Feng Xiong, Liang Lin +5

    cs.CVcs.AIcs.CLarXiv:2604.09349v22026
  57. Scale-Aware Modulation Meet Transformer

    Weifeng Lin, Ziheng Wu, Jiayu Chen +2

    cs.CVarXiv:2307.08579v22023
  58. E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models

    Shengjun Zhang, Zhang Zhang, Chensheng Dai +1

    cs.LGcs.AIcs.CVarXiv:2601.00423v12026
  59. Pixel-level Encoding and Depth Layering for Instance-level Semantic Labeling

    Jonas Uhrig, Marius Cordts, Uwe Franke +1

    cs.CVarXiv:1604.05096v22016
  60. Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation

    Taekyung Ki, Sangwon Jang, Jaehyeong Jo +2

    cs.LGcs.AIcs.CVarXiv:2601.00664v22026