Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,861 to 7,920 of 18,867

  1. OctSqueeze: Octree-Structured Entropy Model for LiDAR Compression

    Lila Huang, Shenlong Wang, Kelvin Wong +2

    eess.IVcs.CVarXiv:2005.07178v22020
  2. Unsupervised Learning of Object Keypoints for Perception and Control

    Tejas Kulkarni, Ankush Gupta, Catalin Ionescu +4

    cs.CVcs.LGarXiv:1906.11883v22019
  3. Blind2Unblind: Self-Supervised Image Denoising with Visible Blind Spots

    Zejin Wang, Jiazheng Liu, Guoqing Li +1

    eess.IVcs.CVarXiv:2203.06967v32022
  4. Co-Mixup: Saliency Guided Joint Mixup with Supermodular Diversity

    Jang-Hyun Kim, Wonho Choo, Hosan Jeong +1

    cs.LGcs.AIcs.CVarXiv:2102.03065v12021
  5. Unsupervised 3D Pose Estimation with Geometric Self-Supervision

    Ching-Hang Chen, Ambrish Tyagi, Amit Agrawal +4

    cs.CVarXiv:1904.04812v12019
  6. Making LLaMA SEE and Draw with SEED Tokenizer

    Yuying Ge, Sijie Zhao, Ziyun Zeng +4

    cs.CVarXiv:2310.01218v12023
  7. A Survey on Content-Aware Video Analysis for Sports

    Huang-Chia Shih

    cs.CVcs.MMarXiv:1703.01170v12017
  8. Confidence-Aware Ensemble and Long-Word Refinement for Artistic Text Recognition

    Lucas A. Dias, Henrique A. Schulz, Rafaela de Miranda +3

    cs.CVarXiv:2608.29970v12026
  9. On the Role of MRI Sequences in Cross-Dataset Generalization for Brain Tumor Segmentation

    Henrique Zan Grande, João G. Pitol, Lucas B. Schuck +3

    cs.CVarXiv:2608.29944v12026
  10. Dynamic Instance Normalization for Arbitrary Style Transfer

    Yongcheng Jing, Xiao Liu, Yukang Ding +4

    cs.CVarXiv:1911.06953v12019
  11. An end-to-end generative framework for video segmentation and recognition

    Hilde Kuehne, Juergen Gall, Thomas Serre

    cs.CVarXiv:1509.01947v22015
  12. AMT: All-Pairs Multi-Field Transforms for Efficient Frame Interpolation

    Zhen Li, Zuo-Liang Zhu, Ling-Hao Han +3

    cs.CVarXiv:2304.09790v12023
  13. CRAFT: Concept Recursive Activation FacTorization for Explainability

    Thomas Fel, Agustin Picard, Louis Bethune +5

    cs.CVcs.AIarXiv:2211.10154v22022
  14. RED: Reinforced Encoder-Decoder Networks for Action Anticipation

    Jiyang Gao, Zhenheng Yang, Ram Nevatia

    cs.CVarXiv:1707.04818v12017
  15. FuncRoom-Agent: Sequential Feed-Forward 3D Functional Indoor Scene Generation

    Hao Feng, Zhi Zuo, MingJian Liang +6

    cs.CVarXiv:2608.29519v12026
  16. When will you do what? - Anticipating Temporal Occurrences of Activities

    Yazan Abu Farha, Alexander Richard, Juergen Gall

    cs.CVarXiv:1804.00892v12018
  17. Seeing Through Extreme Visual Sparsity: Surface Understanding from a Single Random Visual Patch

    Sindhuja Penchala, Sudip Mittal, Noorbakhsh Amiri Golilarz

    cs.CVarXiv:2608.29475v12026
  18. CERF: Communication-Efficient and Retraining-Free Collaborative Perception

    Jiuwu Hao, Ziyi Ni, Liguo Sun +5

    cs.CVarXiv:2609.00951v12026
  19. SG-Net: Spatial Granularity Network for One-Stage Video Instance Segmentation

    Dongfang Liu, Yiming Cui, Wenbo Tan +1

    cs.CVarXiv:2103.10284v22021
  20. Deep CT to MR Synthesis using Paired and Unpaired Data

    Cheng-Bin Jin, Hakil Kim, Wonmo Jung +6

    cs.CVarXiv:1805.10790v22018
  21. Spectral band selection for vegetation properties retrieval using Gaussian processes regression

    Jochem Verrelst, Juan Pablo Rivera, Anatoly Gitelson +3

    cs.CVeess.IVstat.AParXiv:2012.08640v12020
  22. Learning Human Pose Estimation Features with Convolutional Networks

    Arjun Jain, Jonathan Tompson, Mykhaylo Andriluka +2

    cs.CVcs.LGcs.NEarXiv:1312.7302v62013
  23. Computed Tomography Reconstruction Using Deep Image Prior and Learned Reconstruction Methods

    Daniel Otero Baguer, Johannes Leuschner, Maximilian Schmidt

    eess.IVcs.CVcs.LGarXiv:2003.04989v22020
  24. Improving Position Encoding of Transformers for Multivariate Time Series Classification

    Navid Mohammadi Foumani, Chang Wei Tan, Geoffrey I. Webb +1

    cs.LGcs.CVarXiv:2305.16642v12023
  25. VCAR: Training-Free 3DGS Segmentation via View Completeness and Axis-Aware Boundary Refinement

    Kun Cao, Di Wang, Haibin Zhu +4

    cs.CVarXiv:2608.30870v12026
  26. Analytic Dynamics: Learning Physics-Grounded Representation for Fast Intrinsic Dynamics Inference from Monocular Videos

    Jailing Lin, Jikuan Zhang, Jianhua Sun

    cs.CVarXiv:2608.31025v12026
  27. Defense against Universal Adversarial Perturbations

    Naveed Akhtar, Jian Liu, Ajmal Mian

    cs.CVarXiv:1711.05929v32017
  28. Task-Driven Super Resolution: Object Detection in Low-resolution Images

    Muhammad Haris, Greg Shakhnarovich, Norimichi Ukita

    cs.CVarXiv:1803.11316v12018
  29. CDMamba: Incorporating Local Clues into Mamba for Remote Sensing Image Binary Change Detection

    Haotian Zhang, Keyan Chen, Chenyang Liu +3

    cs.CVarXiv:2406.04207v22024
  30. Modality Distillation with Multiple Stream Networks for Action Recognition

    Nuno Garcia, Pietro Morerio, Vittorio Murino

    cs.CVarXiv:1806.07110v22018
  31. Automatic calcium scoring in low-dose chest CT using deep neural networks with dilated convolutions

    Nikolas Lessmann, Bram van Ginneken, Majd Zreik +4

    cs.CVarXiv:1711.00349v22017
  32. Dense Optical Flow Prediction from a Static Image

    Jacob Walker, Abhinav Gupta, Martial Hebert

    cs.CVarXiv:1505.00295v22015
  33. Improving Referring Expression Grounding with Cross-modal Attention-guided Erasing

    Xihui Liu, Zihao Wang, Jing Shao +2

    cs.CVcs.CLarXiv:1903.00839v22019
  34. Music Gesture for Visual Sound Separation

    Chuang Gan, Deng Huang, Hang Zhao +2

    cs.CVcs.LGcs.MMarXiv:2004.09476v12020
  35. Learning Monocular Depth by Distilling Cross-domain Stereo Networks

    Xiaoyang Guo, Hongsheng Li, Shuai Yi +2

    cs.CVarXiv:1808.06586v12018
  36. Anatomy of Catastrophic Forgetting: Hidden Representations and Task Semantics

    Vinay V. Ramasesh, Ethan Dyer, Maithra Raghu

    cs.LGcs.CVstat.MLarXiv:2007.07400v12020
  37. Hierarchical Dynamic Filtering Network for RGB-D Salient Object Detection

    Youwei Pang, Lihe Zhang, Xiaoqi Zhao +1

    cs.CVarXiv:2007.06227v32020
  38. A Composition-Aware Pretraining Framework for Geospatial Foundation Models

    Aryan Kashyap Naveen, Abhishek Srinivas, Pranav Moothedath +1

    cs.CVcs.AIarXiv:2608.30817v12026
  39. Offboard 3D Object Detection from Point Cloud Sequences

    Charles R. Qi, Yin Zhou, Mahyar Najibi +4

    cs.CVarXiv:2103.05073v12021
  40. Evaluation of CNN-based Single-Image Depth Estimation Methods

    Tobias Koch, Lukas Liebel, Friedrich Fraundorfer +1

    cs.CVarXiv:1805.01328v12018
  41. Radiology Report Generation with a Learned Knowledge Base and Multi-modal Alignment

    Shuxin Yang, Xian Wu, Shen Ge +2

    eess.IVcs.CLcs.CVarXiv:2112.15011v22021
  42. Reading Car License Plates Using Deep Convolutional Neural Networks and LSTMs

    Hui Li, Chunhua Shen

    cs.CVarXiv:1601.05610v12016
  43. On Adversarial Robustness of Trajectory Prediction for Autonomous Vehicles

    Qingzhao Zhang, Shengtuo Hu, Jiachen Sun +2

    cs.CVcs.CRarXiv:2201.05057v32022
  44. FusionNet: 3D Object Classification Using Multiple Data Representations

    Vishakh Hegde, Reza Zadeh

    cs.CVarXiv:1607.05695v42016
  45. Dense Pose Transfer

    Natalia Neverova, Riza Alp Guler, Iasonas Kokkinos

    cs.CVarXiv:1809.01995v12018
  46. Real-Time Anomaly Detection and Localization in Crowded Scenes

    Mohammad Sabokrou, Mahmood Fathy, Mojtaba Hosseini +1

    cs.CVarXiv:1511.06936v12015
  47. Models Matter, So Does Training: An Empirical Study of CNNs for Optical Flow Estimation

    Deqing Sun, Xiaodong Yang, Ming-Yu Liu +1

    cs.CVarXiv:1809.05571v12018
  48. FANet: A Feedback Attention Network for Improved Biomedical Image Segmentation

    Nikhil Kumar Tomar, Debesh Jha, Michael A. Riegler +5

    cs.CVeess.IVarXiv:2103.17235v32021
  49. CORAL: A Benchmark for Structure-aware and Brain-wide Neuron Reconstruction in Light Microscopy

    Zekang Yang, Jiamin Li, Zhenghua Li +3

    cs.CVarXiv:2608.30768v12026
  50. Where are the Blobs: Counting by Localization with Point Supervision

    Issam H. Laradji, Negar Rostamzadeh, Pedro O. Pinheiro +2

    cs.CVarXiv:1807.09856v12018
  51. HorizonNet for visual terrain navigation

    Bertil Grelsson, Andreas Robinson, Michael Felsberg +1

    cs.CVcs.ROarXiv:2608.30471v12026
  52. Residual Attention U-Net for Automated Multi-Class Segmentation of COVID-19 Chest CT Images

    Xiaocong Chen, Lina Yao, Yu Zhang

    eess.IVcs.CVcs.LGarXiv:2004.05645v12020
  53. SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model

    Yang Zhan, Zhitong Xiong, Yuan Yuan

    cs.CVarXiv:2401.09712v12024
  54. Motion-Saliency Complementary Masked Modeling for Point Cloud Video Understanding

    Wei Wang, Yiding Sun, Yuyan Wang +4

    cs.CVarXiv:2608.30279v12026
  55. Ground Truth Evaluation of Neural Network Explanations with CLEVR-XAI

    Leila Arras, Ahmed Osman, Wojciech Samek

    cs.CVcs.LGcs.NEarXiv:2003.07258v22020
  56. Visual Dexterity: In-Hand Reorientation of Novel and Complex Object Shapes

    Tao Chen, Megha Tippur, Siyang Wu +3

    cs.ROcs.AIcs.CVarXiv:2211.11744v32022
  57. Cube Padding for Weakly-Supervised Saliency Prediction in 360° Videos

    Hsien-Tzu Cheng, Chun-Hung Chao, Jin-Dong Dong +3

    cs.CVarXiv:1806.01320v12018
  58. Transcending Forgery Specificity with Latent Space Augmentation for Generalizable Deepfake Detection

    Zhiyuan Yan, Yuhao Luo, Siwei Lyu +2

    cs.CVarXiv:2311.11278v22023
  59. Multi-organ Segmentation over Partially Labeled Datasets with Multi-scale Feature Abstraction

    Xi Fang, Pingkun Yan

    cs.CVarXiv:2001.00208v22020
  60. PyMAF-X: Towards Well-aligned Full-body Model Regression from Monocular Images

    Hongwen Zhang, Yating Tian, Yuxiang Zhang +4

    cs.CVarXiv:2207.06400v42022