Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
3,601 to 3,660 of 18,796
Fast Resampling of 3D Point Clouds via Graphs
Siheng Chen, Dong Tian, Chen Feng +2
cs.CVarXiv:1702.06397v12017Bottom Up Top Down Detection Transformers for Language Grounding in Images and Point Clouds
Ayush Jain, Nikolaos Gkanatsios, Ishita Mediratta +1
cs.CVcs.CLarXiv:2112.08879v52021Count-ception: Counting by Fully Convolutional Redundant Counting
Joseph Paul Cohen, Genevieve Boucher, Craig A. Glastonbury +2
cs.CVcs.LGstat.MLarXiv:1703.08710v22017Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
Hao Yang, Zhiyu Tan, Jia Gong +7
cs.CVarXiv:2602.08820v22026Spatial Pyramid Based Graph Reasoning for Semantic Segmentation
Xia Li, Yibo Yang, Qijie Zhao +3
cs.CVarXiv:2003.10211v12020L2G: A Simple Local-to-Global Knowledge Transfer Framework for Weakly Supervised Semantic Segmentation
Peng-Tao Jiang, Yuqi Yang, Qibin Hou +1
cs.CVarXiv:2204.03206v12022Physics-based Noise Modeling for Extreme Low-light Photography
Kaixuan Wei, Ying Fu, Yinqiang Zheng +1
eess.IVcs.CVarXiv:2108.02158v12021MIRROR: Manifold Ideal Reference ReconstructOR for Generalizable AI-Generated Image Detection
Ruiqi Liu, Manni Cui, Ziheng Qin +12
cs.CVcs.CRarXiv:2602.02222v12026Provable Dynamic Fusion for Low-Quality Multimodal Data
Qingyang Zhang, Haitao Wu, Changqing Zhang +4
cs.LGcs.CVarXiv:2306.02050v22023MV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D Generation
Baicheng Li, Dong Wu, Jun Li +4
cs.CVarXiv:2603.11633v22026Clustering Millions of Faces by Identity
Charles Otto, Dayong Wang, Anil K. Jain
cs.CVarXiv:1604.00989v12016Zero-Shot Grounding of Objects from Natural Language Queries
Arka Sadhu, Kan Chen, Ram Nevatia
cs.CVcs.CLarXiv:1908.07129v12019Fast Low-rank Shared Dictionary Learning for Image Classification
Tiep Vu, Vishal Monga
cs.CVcs.AIarXiv:1610.08606v32016Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models
Hao Tang, Yu Liu, Shuanglin Yan +3
cs.CVcs.MMarXiv:2601.08476v22026Contour-Guided Query-Based Feature Fusion for Boundary-Aware and Generalizable Cardiac Ultrasound Segmentation
Zahid Ullah, Sieun Choi, Jihie Kim
cs.CVarXiv:2603.28110v12026ReRoPE: Repurposing RoPE for Relative Camera Control
Chunyang Li, Yuanbo Yang, Jiahao Shao +3
cs.CVarXiv:2602.08068v12026Learning Fast, Learning Slow: A General Continual Learning Method based on Complementary Learning System
Elahe Arani, Fahad Sarfraz, Bahram Zonooz
cs.LGcs.AIcs.CVarXiv:2201.12604v22022DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving
Shuyao Shang, Bing Zhan, Yunfei Yan +9
cs.CVcs.ROarXiv:2603.11041v22026Diverse Weight Averaging for Out-of-Distribution Generalization
Alexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier +3
cs.CVcs.AIcs.LGarXiv:2205.09739v22022X-View: Graph-Based Semantic Multi-View Localization
Abel Gawel, Carlo Del Don, Roland Siegwart +2
cs.ROcs.CVarXiv:1709.09905v32017TC-SSA: Token Compression via Semantic Slot Aggregation for Gigapixel Pathology Reasoning
Zhuo Chen, Shawn Young, Lijian Xu
cs.CVcs.AIarXiv:2603.01143v12026Fully Convolutional Crowd Counting On Highly Congested Scenes
Mark Marsden, Kevin McGuinness, Suzanne Little +1
cs.CVarXiv:1612.00220v22016Lifelong Machine Learning with Deep Streaming Linear Discriminant Analysis
Tyler L. Hayes, Christopher Kanan
cs.LGcs.CVstat.MLarXiv:1909.01520v32019Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models
Zhiqiu Lin, Samuel Yu, Zhiyi Kuang +2
cs.CVcs.AIcs.LGarXiv:2301.06267v52023When NAS Meets Robustness: In Search of Robust Architectures against Adversarial Attacks
Minghao Guo, Yuzhe Yang, Rui Xu +2
cs.LGcs.CRcs.CVarXiv:1911.10695v32019Trans-SVNet: Accurate Phase Recognition from Surgical Videos via Hybrid Embedding Aggregation Transformer
Xiaojie Gao, Yueming Jin, Yonghao Long +2
cs.CVcs.AIarXiv:2103.09712v22021TransCenter: Transformers with Dense Representations for Multiple-Object Tracking
Yihong Xu, Yutong Ban, Guillaume Delorme +3
cs.CVarXiv:2103.15145v42021HF-NeuS: Improved Surface Reconstruction Using High-Frequency Details
Yiqun Wang, Ivan Skorokhodov, Peter Wonka
cs.CVcs.GRarXiv:2206.07850v22022Visual Question Generation as Dual Task of Visual Question Answering
Yikang Li, Nan Duan, Bolei Zhou +3
cs.CVarXiv:1709.07192v12017SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
Jiwen Zhang, Zejun Li, Siyuan Wang +3
cs.CVcs.AIcs.ROarXiv:2601.06806v12026Context-Aware Trajectory Prediction
Federico Bartoli, Giuseppe Lisanti, Lamberto Ballan +1
cs.CVarXiv:1705.02503v12017TF-Blender: Temporal Feature Blender for Video Object Detection
Yiming Cui, Liqi Yan, Zhiwen Cao +1
cs.CVarXiv:2108.05821v12021ARM: Advantage Reward Modeling for Long-Horizon Manipulation
Yiming Mao, Zixi Yu, Weixin Mao +5
cs.ROcs.AIcs.CVarXiv:2604.03037v22026NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
Huichao Zhang, Liao Qu, Yiheng Liu +33
cs.CVcs.AIarXiv:2601.02204v120263D Human Pose Estimation in RGBD Images for Robotic Task Learning
Christian Zimmermann, Tim Welschehold, Christian Dornhege +2
cs.CVcs.ROarXiv:1803.02622v22018AutoFigure-Edit: Generating Editable Scientific Illustration
Zhen Lin, Qiujie Xie, Minjun Zhu +10
cs.CVcs.AIarXiv:2603.06674v12026Synthetic data generation for end-to-end thermal infrared tracking
Lichao Zhang, Abel Gonzalez-Garcia, Joost van de Weijer +2
cs.CVarXiv:1806.01013v22018Landslide4Sense: Reference Benchmark Data and Deep Learning Models for Landslide Detection
Omid Ghorbanzadeh, Yonghao Xu, Pedram Ghamisi +2
cs.CVeess.IVarXiv:2206.00515v32022SynQ: Accurate Zero-shot Quantization by Synthesis-aware Fine-tuning
Minjun Kim, Jongjin Kim, U Kang
cs.CVarXiv:2603.18423v12026Prostate Cancer Diagnosis using Deep Learning with 3D Multiparametric MRI
Saifeng Liu, Huaixiu Zheng, Yesu Feng +1
cs.CVstat.MLarXiv:1703.04078v12017Hierarchical Dense Correlation Distillation for Few-Shot Segmentation
Bohao Peng, Zhuotao Tian, Xiaoyang Wu +4
cs.CVcs.AIarXiv:2303.14652v12023AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots
Likui Zhang, Tao Tang, Zhihao Zhan +9
cs.ROcs.AIcs.CVarXiv:2603.07648v12026Self-Supervised Learning of Audio-Visual Objects from Video
Triantafyllos Afouras, Andrew Owens, Joon Son Chung +1
cs.CVcs.SDeess.ASarXiv:2008.04237v12020Online Video Deblurring via Dynamic Temporal Blending Network
Tae Hyun Kim, Kyoung Mu Lee, Bernhard Schölkopf +1
cs.CVarXiv:1704.03285v12017Full-Duplex Strategy for Video Object Segmentation
Ge-Peng Ji, Deng-Ping Fan, Keren Fu +3
cs.CVarXiv:2108.03151v32021LoMa: Local Feature Matching Revisited
David Nordström, Johan Edstedt, Georg Bökman +6
cs.CVarXiv:2604.04931v22026Deep Gaussian Scale Mixture Prior for Spectral Compressive Imaging
Tao Huang, Weisheng Dong, Xin Yuan +2
eess.IVcs.CVarXiv:2103.07152v22021Point Cloud Upsampling via Disentangled Refinement
Ruihui Li, Xianzhi Li, Pheng-Ann Heng +1
cs.CVarXiv:2106.04779v12021TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
Chuancheng Shi, Shangze Li, Wenjun Lu +5
cs.CVcs.AIcs.CYarXiv:2601.21900v22026Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting
Hao Feng, Wei Shi, Ke Zhang +9
cs.CVarXiv:2602.05384v12026Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
Lili Yu, Bowen Shi, Ramakanth Pasunuru +24
cs.LGcs.CLcs.CVarXiv:2309.02591v12023PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction
Peng Wang, Hao Tan, Sai Bi +6
cs.CVarXiv:2311.12024v22023Deep DIC: Deep Learning-Based Digital Image Correlation for End-to-End Displacement and Strain Measurement
Ru Yang, Yang Li, Danielle Zeng +1
eess.IVcond-mat.mtrl-scics.CVarXiv:2110.13720v22021Contrastive and Non-Contrastive Self-Supervised Learning Recover Global and Local Spectral Embedding Methods
Randall Balestriero, Yann LeCun
cs.LGcs.AIcs.CVarXiv:2205.11508v32022RenderGAN: Generating Realistic Labeled Data
Leon Sixt, Benjamin Wild, Tim Landgraf
cs.NEcs.CVarXiv:1611.01331v52016Visually-Guided Policy Optimization for Multimodal Reasoning
Zengbin Wang, Feng Xiong, Liang Lin +5
cs.CVcs.AIcs.CLarXiv:2604.09349v22026Scale-Aware Modulation Meet Transformer
Weifeng Lin, Ziheng Wu, Jiayu Chen +2
cs.CVarXiv:2307.08579v22023E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models
Shengjun Zhang, Zhang Zhang, Chensheng Dai +1
cs.LGcs.AIcs.CVarXiv:2601.00423v12026Pixel-level Encoding and Depth Layering for Instance-level Semantic Labeling
Jonas Uhrig, Marius Cordts, Uwe Franke +1
cs.CVarXiv:1604.05096v22016Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
Taekyung Ki, Sangwon Jang, Jaehyeong Jo +2
cs.LGcs.AIcs.CVarXiv:2601.00664v22026