Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

9,301 to 9,360 of 18,811

  1. How Far Can 5,500 Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models

    Victor Besnier, Anh-Quan Cao, Elias Ramzi +5

    cs.CVarXiv:2608.28404v12026
  2. GradNet: Gradient-Guided Network for Visual Object Tracking

    Peixia Li, Boyu Chen, Wanli Ouyang +3

    cs.CVarXiv:1909.06800v12019
  3. ImVoxelNet: Image to Voxels Projection for Monocular and Multi-View General-Purpose 3D Object Detection

    Danila Rukhovich, Anna Vorontsova, Anton Konushin

    cs.CVarXiv:2106.01178v32021
  4. Brain Network Transformer

    Xuan Kan, Wei Dai, Hejie Cui +3

    cs.LGcs.CVcs.NEarXiv:2210.06681v22022
  5. Conditional Visual Evidence Utility: State-Dependent Rank Reversals in Frozen Vision-Language Encoders

    Yunxuan Fang, Xinhe Wang

    cs.CVarXiv:2608.28316v12026
  6. Towards Robust Vision Transformer

    Xiaofeng Mao, Gege Qi, Yuefeng Chen +5

    cs.CVarXiv:2105.07926v42021
  7. Temporal Context Network for Activity Localization in Videos

    Xiyang Dai, Bharat Singh, Guyue Zhang +2

    cs.CVarXiv:1708.02349v12017
  8. Prompt-Guided Interactive Segmentation of Interstitial Lung Disease in Thoracic CT

    Vasilis Dedousis, Lubnaa Abdur Rahman, Lorenzo Brigatο +8

    cs.CVarXiv:2608.28453v12026
  9. CycleGAN, a Master of Steganography

    Casey Chu, Andrey Zhmoginov, Mark Sandler

    cs.CVcs.LGstat.MLarXiv:1712.02950v22017
  10. μ-MAR: Multiplane 3D Marker based Registration for Depth-sensing Cameras

    Marcelo Saval-Calvo, Jorge Azorin-Lopez, Andres Fuster-Guillo +1

    cs.CVarXiv:1708.01405v12017
  11. Depth-Regularized Optimization for 3D Gaussian Splatting in Few-Shot Images

    Jaeyoung Chung, Jeongtaek Oh, Kyoung Mu Lee

    cs.CVcs.GRarXiv:2311.13398v32023
  12. VSGNet: Spatial Attention Network for Detecting Human Object Interactions Using Graph Convolutions

    Oytun Ulutan, A S M Iftekhar, B. S. Manjunath

    cs.CVarXiv:2003.05541v12020
  13. Denoising-Aware Temporal Point Cloud Completion for 3D Crop Architecture Recovery and Phenotypic Trait Extraction

    Mrudul Mittal, Soumyashree Kar

    cs.CVarXiv:2608.28343v12026
  14. MonoDTR: Monocular 3D Object Detection with Depth-Aware Transformer

    Kuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su +1

    cs.CVarXiv:2203.10981v22022
  15. REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction

    Zeyi Liu, Arpit Bahety, Shuran Song

    cs.ROcs.AIcs.CLarXiv:2306.15724v42023
  16. LightTrack: Finding Lightweight Neural Networks for Object Tracking via One-Shot Architecture Search

    Bin Yan, Houwen Peng, Kan Wu +3

    cs.CVarXiv:2104.14545v12021
  17. High-resolution Iterative Feedback Network for Camouflaged Object Detection

    Xiaobin Hu, Shuo Wang, Xuebin Qin +5

    cs.CVarXiv:2203.11624v22022
  18. Abstract4D: A Large-Scale Dataset and Framework for Understanding the Visual Language of Abstract Art

    Haowei Zhang, Yuanpei Zhao, Ji-Zhe Zhou +1

    cs.CVarXiv:2608.28339v12026
  19. Predicting Deeper into the Future of Semantic Segmentation

    Pauline Luc, Natalia Neverova, Camille Couprie +2

    cs.CVcs.LGarXiv:1703.07684v32017
  20. Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like Speed

    Yifan Wang, Xingyi He, Sida Peng +2

    cs.CVarXiv:2403.04765v22024
  21. Parameter-free Online Test-time Adaptation

    Malik Boudiaf, Romain Mueller, Ismail Ben Ayed +1

    cs.CVarXiv:2201.05718v22022
  22. Diff-Instruct: A Universal Approach for Transferring Knowledge From Pre-trained Diffusion Models

    Weijian Luo, Tianyang Hu, Shifeng Zhang +3

    cs.LGcs.CVarXiv:2305.18455v22023
  23. FUSED: Forensic-Semantic Mixture-of-Experts for AI Inpainting Detection and Localization

    Anton Nuzhdin, Marcel Worring, Ivona Najdenkoska

    cs.CVarXiv:2608.28302v12026
  24. GeoFF3D: Coordinate-Anchored Feed-Forward Reconstruction for Large-Scale UAV Mapping

    Xiang Yang, Yongli Wang, Yunsheng Zhang

    cs.CVarXiv:2608.28288v12026
  25. DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning

    Yujie Wei, Xinyu Liu, Shiwei Zhang +12

    cs.CVarXiv:2603.12257v12026
  26. SparseNeRF: Distilling Depth Ranking for Few-shot Novel View Synthesis

    Guangcong Wang, Zhaoxi Chen, Chen Change Loy +1

    cs.CVarXiv:2303.16196v22023
  27. Transformer for Image Quality Assessment

    Junyong You, Jari Korhonen

    cs.CVcs.LGeess.IVarXiv:2101.01097v22020
  28. Enhanced discrete particle swarm optimization path planning for UAV vision-based surface inspection

    Manh Duong Phung, Cong Hoang Quach, Tran Hiep Dinh +1

    cs.ROcs.AIcs.CVarXiv:1706.04399v12017
  29. Cut-ViT: Task-Specific Model Pruning via Gram Anchoring Subspace Consistency

    Jianjian Yin, Liulei Li, Tao Chen +3

    cs.CVarXiv:2608.28205v12026
  30. Cyc3D: Evaluating Cyclic Structural Stability and Asset Usability in Image-to-3D Generation

    Liwen Zhang

    cs.CVarXiv:2608.28080v12026
  31. VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

    Tsu-Jui Fu, Linjie Li, Zhe Gan +4

    cs.CVarXiv:2111.12681v22021
  32. Non-Uniform Quantisation for 3DGS Compression

    Bert Van hauwermeiren, Patrice Rondao Alface, Adrian Munteanu

    cs.CVarXiv:2608.28272v12026
  33. Context Based Emotion Recognition using EMOTIC Dataset

    Ronak Kosti, Jose M. Alvarez, Adria Recasens +1

    cs.CVcs.LGarXiv:2003.13401v12020
  34. Virchow2: Scaling Self-Supervised Mixed Magnification Models in Pathology

    Eric Zimmermann, Eugene Vorontsov, Julian Viret +11

    cs.CVarXiv:2408.00738v32024
  35. WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild

    Yuhao Bai, Qianqiu Tan, Lilong Chen +2

    cs.CVarXiv:2608.28240v12026
  36. Learning Diverse and Discriminative Representations via the Principle of Maximal Coding Rate Reduction

    Yaodong Yu, Kwan Ho Ryan Chan, Chong You +2

    cs.LGcs.CVcs.ITarXiv:2006.08558v12020
  37. RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation

    Zhen Xiao, Zhen Shen, Zhaofan Qiu +3

    cs.CVarXiv:2608.28219v12026
  38. Diversity in Machine Learning

    Zhiqiang Gong, Ping Zhong, Weidong Hu

    cs.CVarXiv:1807.01477v22018
  39. WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes

    Kishor Datta Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman +2

    cs.CVarXiv:2608.28216v12026
  40. CoCoNet: Coupled Contrastive Learning Network with Multi-level Feature Ensemble for Multi-modality Image Fusion

    Jinyuan Liu, Runjia Lin, Guanyao Wu +3

    cs.CVarXiv:2211.10960v32022
  41. ISLES 2022: A multi-center magnetic resonance imaging stroke lesion segmentation dataset

    Moritz Roman Hernandez Petzsche, Ezequiel de la Rosa, Uta Hanning +22

    cs.CVarXiv:2206.06694v12022
  42. Boosting R-CNN: Reweighting R-CNN Samples by RPN's Error for Underwater Object Detection

    Pinhao Song, Pengteng Li, Linhui Dai +2

    cs.CVarXiv:2206.13728v32022
  43. Deep learning based cloud detection for medium and high resolution remote sensing images of different sensors

    Zhiwei Li, Huanfeng Shen, Qing Cheng +3

    cs.CVarXiv:1810.05801v32018
  44. Adversarial Feature Augmentation for Unsupervised Domain Adaptation

    Riccardo Volpi, Pietro Morerio, Silvio Savarese +1

    cs.CVarXiv:1711.08561v22017
  45. Conditional Image-to-Video Generation with Latent Flow Diffusion Models

    Haomiao Ni, Changhao Shi, Kai Li +2

    cs.CVarXiv:2303.13744v12023
  46. Post-Training VLMs for Video Mistake Detection

    Federico Spurio, Olga Zatsarynna, Lars Doorenbos +3

    cs.CVcs.LGarXiv:2608.28406v12026
  47. Towards Conceptual Compression

    Karol Gregor, Frederic Besse, Danilo Jimenez Rezende +2

    stat.MLcs.CVcs.LGarXiv:1604.08772v12016
  48. Instance Segmentation by Jointly Optimizing Spatial Embeddings and Clustering Bandwidth

    Davy Neven, Bert De Brabandere, Marc Proesmans +1

    cs.CVarXiv:1906.11109v22019
  49. Enhanced Invertible Encoding for Learned Image Compression

    Yueqi Xie, Ka Leong Cheng, Qifeng Chen

    eess.IVcs.CVarXiv:2108.03690v12021
  50. Bringing Old Photos Back to Life

    Ziyu Wan, Bo Zhang, Dongdong Chen +4

    cs.CVcs.GReess.IVarXiv:2004.09484v12020
  51. NeST: A Neural Network Synthesis Tool Based on a Grow-and-Prune Paradigm

    Xiaoliang Dai, Hongxu Yin, Niraj K. Jha

    cs.NEcs.AIcs.CVarXiv:1711.02017v32017
  52. Spherical Transformer for LiDAR-based 3D Recognition

    Xin Lai, Yukang Chen, Fanbin Lu +2

    cs.CVcs.AIarXiv:2303.12766v12023
  53. Learning Robust Representations by Projecting Superficial Statistics Out

    Haohan Wang, Zexue He, Zachary C. Lipton +1

    cs.CVcs.LGarXiv:1903.06256v12019
  54. LCDNet: Deep Loop Closure Detection and Point Cloud Registration for LiDAR SLAM

    Daniele Cattaneo, Matteo Vaghi, Abhinav Valada

    cs.ROcs.CVcs.LGarXiv:2103.05056v42021
  55. A Controlled Audit of Architectural Complexity in Uncertainty-Aware Multi-Organ Ultrasound Classification

    Yang Song, Pengbo Sun, Shichang Feng +3

    cs.CVarXiv:2608.28063v12026
  56. Boundary-preserving Mask R-CNN

    Tianheng Cheng, Xinggang Wang, Lichao Huang +1

    cs.CVarXiv:2007.08921v12020
  57. Joint Detection and Recounting of Abnormal Events by Learning Deep Generic Knowledge

    Ryota Hinami, Tao Mei, Shin'ichi Satoh

    cs.CVarXiv:1709.09121v12017
  58. Gaussian-SLAM: Photo-realistic Dense SLAM with Gaussian Splatting

    Vladimir Yugay, Yue Li, Theo Gevers +1

    cs.CVcs.ROarXiv:2312.10070v22023
  59. Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and Localization

    Komal Chugh, Parul Gupta, Abhinav Dhall +1

    cs.CVcs.MMarXiv:2005.14405v32020
  60. PillarNet: Real-Time and High-Performance Pillar-based 3D Object Detection

    Guangsheng Shi, Ruifeng Li, Chao Ma

    cs.CVarXiv:2205.07403v52022