Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,741 to 13,800 of 18,802

  1. A Survey on Deep Neural Network Pruning-Taxonomy, Comparison, Analysis, and Recommendations

    Hongrong Cheng, Miao Zhang, Javen Qinfeng Shi

    cs.LGcs.CVarXiv:2308.06767v22023
  2. Mapping the stereotyped behaviour of freely-moving fruit flies

    Gordon J. Berman, Daniel M. Choi, William Bialek +1

    q-bio.QMcs.CVphysics.bio-pharXiv:1310.4249v22013
  3. ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning

    Qiao Gu, Alihusein Kuwajerwala, Sacha Morin +13

    cs.ROcs.CVarXiv:2309.16650v12023
  4. Graph Convolutional Label Noise Cleaner: Train a Plug-and-play Action Classifier for Anomaly Detection

    Jia-Xing Zhong, Nannan Li, Weijie Kong +3

    cs.CVarXiv:1903.07256v12019
  5. Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes

    Martin Sundermeyer, Arsalan Mousavian, Rudolph Triebel +1

    cs.ROcs.CVarXiv:2103.14127v12021
  6. PANDA: Pose Aligned Networks for Deep Attribute Modeling

    Ning Zhang, Manohar Paluri, Marc'Aurelio Ranzato +2

    cs.CVarXiv:1311.5591v22013
  7. A Unified Model for Multi-class Anomaly Detection

    Zhiyuan You, Lei Cui, Yujun Shen +4

    cs.CVarXiv:2206.03687v32022
  8. Long Context Transfer from Language to Vision

    Peiyuan Zhang, Kaichen Zhang, Bo Li +7

    cs.CVarXiv:2406.16852v22024
  9. Drone-based Object Counting by Spatially Regularized Regional Proposal Network

    Meng-Ru Hsieh, Yen-Liang Lin, Winston H. Hsu

    cs.CVarXiv:1707.05972v32017
  10. Spikformer: When Spiking Neural Network Meets Transformer

    Zhaokun Zhou, Yuesheng Zhu, Chao He +4

    cs.NEcs.CVcs.LGarXiv:2209.15425v22022
  11. 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and Camera

    Iro Armeni, Zhi-Yang He, JunYoung Gwak +4

    cs.CVcs.ROarXiv:1910.02527v12019
  12. An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA

    Zhengyuan Yang, Zhe Gan, Jianfeng Wang +4

    cs.CVarXiv:2109.05014v22021
  13. Human pose estimation via Convolutional Part Heatmap Regression

    Adrian Bulat, Georgios Tzimiropoulos

    cs.CVarXiv:1609.01743v12016
  14. Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition

    Pengfei Zhang, Cuiling Lan, Wenjun Zeng +3

    cs.CVarXiv:1904.01189v32019
  15. DeepJDOT: Deep Joint Distribution Optimal Transport for Unsupervised Domain Adaptation

    Bharath Bhushan Damodaran, Benjamin Kellenberger, Rémi Flamary +2

    cs.CVcs.AIarXiv:1803.10081v32018
  16. The 2018 PIRM Challenge on Perceptual Image Super-resolution

    Yochai Blau, Roey Mechrez, Radu Timofte +2

    cs.CVarXiv:1809.07517v32018
  17. Easily Accessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale

    Federico Bianchi, Pratyusha Kalluri, Esin Durmus +7

    cs.CLcs.CVarXiv:2211.03759v22022
  18. Skeleton-Based Human Action Recognition with Global Context-Aware Attention LSTM Networks

    Jun Liu, Gang Wang, Ling-Yu Duan +2

    cs.CVarXiv:1707.05740v52017
  19. GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model

    Md Selim Sarowar, Md Tanvir Islam, Sungho Kim +1

    cs.ROcs.CVarXiv:2608.24959v12026
  20. Visual Attribute Transfer through Deep Image Analogy

    Jing Liao, Yuan Yao, Lu Yuan +2

    cs.CVarXiv:1705.01088v22017
  21. VisDocAgentBench: Benchmarking Agents for Visually Rich Document Retrieval

    Lexiang Hu, Yanzhao Zhang, Mingxin Li +5

    cs.IRcs.AIcs.CVarXiv:2608.17889v12026
  22. Total-Text: A Comprehensive Dataset for Scene Text Detection and Recognition

    Chee Kheng Chng, Chee Seng Chan

    cs.CVarXiv:1710.10400v12017
  23. Automatic Brain Tumor Segmentation using Cascaded Anisotropic Convolutional Neural Networks

    Guotai Wang, Wenqi Li, Sebastien Ourselin +1

    cs.CVarXiv:1709.00382v22017
  24. The Perfect Match: 3D Point Cloud Matching with Smoothed Densities

    Zan Gojcic, Caifa Zhou, Jan D. Wegner +1

    cs.CVarXiv:1811.06879v32018
  25. Improving Cross-Site Whole-Heart Segmentation

    Tanish Mudaliar, Justin Li, Daniel Lin +4

    eess.IVcs.CVarXiv:2608.25109v12026
  26. Image-based localization using LSTMs for structured feature correlation

    Florian Walch, Caner Hazirbas, Laura Leal-Taixé +3

    cs.CVarXiv:1611.07890v42016
  27. Three Factors Influencing Minima in SGD

    Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit +4

    cs.LGcs.AIcs.CVarXiv:1711.04623v32017
  28. Knowledge Distillation by On-the-Fly Native Ensemble

    Xu Lan, Xiatian Zhu, Shaogang Gong

    cs.CVarXiv:1806.04606v22018
  29. Spiking-YOLO: Spiking Neural Network for Energy-Efficient Object Detection

    Seijoon Kim, Seongsik Park, Byunggook Na +1

    cs.CVcs.LGstat.MLarXiv:1903.06530v22019
  30. Lightweight Machine Learning-Driven Monocular Sidewalk Path Extraction for Embedded Micromobility Navigation

    Lkhanaajav Mijiddorj, Yang Yan, Tyler Beringer +4

    cs.CVcs.AIarXiv:2608.25178v12026
  31. Tracking The Untrackable: Learning To Track Multiple Cues with Long-Term Dependencies

    Amir Sadeghian, Alexandre Alahi, Silvio Savarese

    cs.CVarXiv:1701.01909v22017
  32. Bounding Box Regression with Uncertainty for Accurate Object Detection

    Yihui He, Chenchen Zhu, Jianren Wang +2

    cs.CVarXiv:1809.08545v32018
  33. ConsensusTAS: Self-Supervised Temporal Action Segmentation for Long-Horizon Construction Videos

    Xiaoshan Zhou, Yafei Sun

    cs.CVarXiv:2608.24043v12026
  34. MDLatLRR: A novel decomposition method for infrared and visible image fusion

    Hui Li, Xiao-Jun Wu, Josef Kittler

    cs.CVarXiv:1811.02291v52018
  35. Can You Trust Frozen Hematology Foundation Models under Acquisition Shift?

    Jai Kumar Sharma, Peeyush Tapadiya

    cs.CVcs.AIq-bio.QMarXiv:2608.25148v12026
  36. SR-LSTM: State Refinement for LSTM towards Pedestrian Trajectory Prediction

    Pu Zhang, Wanli Ouyang, Pengfei Zhang +2

    cs.CVarXiv:1903.02793v12019
  37. Human Action Recognition using Factorized Spatio-Temporal Convolutional Networks

    Lin Sun, Kui Jia, Dit-Yan Yeung +1

    cs.CVarXiv:1510.00562v12015
  38. GLaMM: Pixel Grounding Large Multimodal Model

    Hanoona Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly +7

    cs.CVcs.AIarXiv:2311.03356v32023
  39. Deep Learning for LiDAR Point Clouds in Autonomous Driving: A Review

    Ying Li, Lingfei Ma, Zilong Zhong +4

    cs.CVarXiv:2005.09830v12020
  40. Lowering the Barrier to AI-Driven Inspection: A No-Code Workflow for Automated Structural Defect Detection

    Michael Holm, Tanner McElroy, Xinghang Zhang +1

    cs.CVcs.LGeess.IVarXiv:2608.25176v12026
  41. SynSin: End-to-end View Synthesis from a Single Image

    Olivia Wiles, Georgia Gkioxari, Richard Szeliski +1

    cs.CVarXiv:1912.08804v22019
  42. ActionFormer: Localizing Moments of Actions with Transformers

    Chenlin Zhang, Jianxin Wu, Yin Li

    cs.CVarXiv:2202.07925v22022
  43. Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

    Ranjan Sapkota, Konstantinos I. Roumeliotis, Pengyao Xie +3

    cs.CVcs.AIarXiv:2608.24934v12026
  44. Modality Contribution Score - A Per-Patient Framework for Quantifying the Relative Diagnostic Contribution of Structural MRI and Amyloid PET in Alzheimer's Disease

    Dawa Chyophel Lepcha, Aaliya Ali, Sophie A. Martin +3

    eess.IVcs.AIcs.CVarXiv:2608.24931v12026
  45. What Do Audio-Visual Synchronization Metrics Actually Measure?

    Jai Kumar Sharma, Peeyush Tapadiya

    cs.CVcs.MMcs.SDarXiv:2608.25157v12026
  46. Dynamic View Synthesis from Dynamic Monocular Video

    Chen Gao, Ayush Saraf, Johannes Kopf +1

    cs.CVarXiv:2105.06468v12021
  47. See More, Detect Less? Taming Information Leakage in Multi-View Anomaly Detection

    Shang-Fu Chen, Kuan-Chuan Peng, Jhih-Ciang Wu +2

    cs.CVcs.MMarXiv:2608.25168v12026
  48. Video Based Reconstruction of 3D People Models

    Thiemo Alldieck, Marcus Magnor, Weipeng Xu +2

    cs.CVarXiv:1803.04758v32018
  49. Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More

    Chaofang Ma, Lin Jiang, Carol Jingyi Li +4

    cs.CVarXiv:2608.25332v12026
  50. GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models

    Yiqun Sun, Junyu Chen, Pengfei Wei +1

    cs.CYcs.CLcs.CVarXiv:2608.25375v12026
  51. Space-time Neural Irradiance Fields for Free-Viewpoint Video

    Wenqi Xian, Jia-Bin Huang, Johannes Kopf +1

    cs.CVarXiv:2011.12950v22020
  52. CRIS: CLIP-Driven Referring Image Segmentation

    Zhaoqing Wang, Yu Lu, Qiang Li +4

    cs.CVarXiv:2111.15174v22021
  53. Modulating early visual processing by language

    Harm de Vries, Florian Strub, Jérémie Mary +3

    cs.CVcs.CLcs.LGarXiv:1707.00683v32017
  54. Submanifold Sparse Convolutional Networks

    Benjamin Graham, Laurens van der Maaten

    cs.NEcs.CVarXiv:1706.01307v12017
  55. 3D-SIS: 3D Semantic Instance Segmentation of RGB-D Scans

    Ji Hou, Angela Dai, Matthias Nießner

    cs.CVarXiv:1812.07003v32018
  56. Bird Species Categorization Using Pose Normalized Deep Convolutional Nets

    Steve Branson, Grant Van Horn, Serge Belongie +1

    cs.CVarXiv:1406.2952v12014
  57. Unprocessing Images for Learned Raw Denoising

    Tim Brooks, Ben Mildenhall, Tianfan Xue +3

    cs.CVcs.LGarXiv:1811.11127v12018
  58. Domain Adaptation for Visual Applications: A Comprehensive Survey

    Gabriela Csurka

    cs.CVarXiv:1702.05374v22017
  59. Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Exactly Detection Tasks

    Weimin Zhou

    eess.IVcs.AIcs.CVarXiv:2608.24768v12026
  60. Asymmetric Cross-Modal Fine-Grained Visual Categorization: ACF-Net and the BirdPro Benchmark

    Bohan Deng, Shuo Ye, Zitong Yu

    cs.CVarXiv:2608.25520v12026