Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
7,861 to 7,920 of 18,867
OctSqueeze: Octree-Structured Entropy Model for LiDAR Compression
Lila Huang, Shenlong Wang, Kelvin Wong +2
eess.IVcs.CVarXiv:2005.07178v22020Unsupervised Learning of Object Keypoints for Perception and Control
Tejas Kulkarni, Ankush Gupta, Catalin Ionescu +4
cs.CVcs.LGarXiv:1906.11883v22019Blind2Unblind: Self-Supervised Image Denoising with Visible Blind Spots
Zejin Wang, Jiazheng Liu, Guoqing Li +1
eess.IVcs.CVarXiv:2203.06967v32022Co-Mixup: Saliency Guided Joint Mixup with Supermodular Diversity
Jang-Hyun Kim, Wonho Choo, Hosan Jeong +1
cs.LGcs.AIcs.CVarXiv:2102.03065v12021Unsupervised 3D Pose Estimation with Geometric Self-Supervision
Ching-Hang Chen, Ambrish Tyagi, Amit Agrawal +4
cs.CVarXiv:1904.04812v12019Making LLaMA SEE and Draw with SEED Tokenizer
Yuying Ge, Sijie Zhao, Ziyun Zeng +4
cs.CVarXiv:2310.01218v12023A Survey on Content-Aware Video Analysis for Sports
Huang-Chia Shih
cs.CVcs.MMarXiv:1703.01170v12017Confidence-Aware Ensemble and Long-Word Refinement for Artistic Text Recognition
Lucas A. Dias, Henrique A. Schulz, Rafaela de Miranda +3
cs.CVarXiv:2608.29970v12026On the Role of MRI Sequences in Cross-Dataset Generalization for Brain Tumor Segmentation
Henrique Zan Grande, João G. Pitol, Lucas B. Schuck +3
cs.CVarXiv:2608.29944v12026Dynamic Instance Normalization for Arbitrary Style Transfer
Yongcheng Jing, Xiao Liu, Yukang Ding +4
cs.CVarXiv:1911.06953v12019An end-to-end generative framework for video segmentation and recognition
Hilde Kuehne, Juergen Gall, Thomas Serre
cs.CVarXiv:1509.01947v22015AMT: All-Pairs Multi-Field Transforms for Efficient Frame Interpolation
Zhen Li, Zuo-Liang Zhu, Ling-Hao Han +3
cs.CVarXiv:2304.09790v12023CRAFT: Concept Recursive Activation FacTorization for Explainability
Thomas Fel, Agustin Picard, Louis Bethune +5
cs.CVcs.AIarXiv:2211.10154v22022RED: Reinforced Encoder-Decoder Networks for Action Anticipation
Jiyang Gao, Zhenheng Yang, Ram Nevatia
cs.CVarXiv:1707.04818v12017FuncRoom-Agent: Sequential Feed-Forward 3D Functional Indoor Scene Generation
Hao Feng, Zhi Zuo, MingJian Liang +6
cs.CVarXiv:2608.29519v12026When will you do what? - Anticipating Temporal Occurrences of Activities
Yazan Abu Farha, Alexander Richard, Juergen Gall
cs.CVarXiv:1804.00892v12018Seeing Through Extreme Visual Sparsity: Surface Understanding from a Single Random Visual Patch
Sindhuja Penchala, Sudip Mittal, Noorbakhsh Amiri Golilarz
cs.CVarXiv:2608.29475v12026CERF: Communication-Efficient and Retraining-Free Collaborative Perception
Jiuwu Hao, Ziyi Ni, Liguo Sun +5
cs.CVarXiv:2609.00951v12026SG-Net: Spatial Granularity Network for One-Stage Video Instance Segmentation
Dongfang Liu, Yiming Cui, Wenbo Tan +1
cs.CVarXiv:2103.10284v22021Deep CT to MR Synthesis using Paired and Unpaired Data
Cheng-Bin Jin, Hakil Kim, Wonmo Jung +6
cs.CVarXiv:1805.10790v22018Spectral band selection for vegetation properties retrieval using Gaussian processes regression
Jochem Verrelst, Juan Pablo Rivera, Anatoly Gitelson +3
cs.CVeess.IVstat.AParXiv:2012.08640v12020Learning Human Pose Estimation Features with Convolutional Networks
Arjun Jain, Jonathan Tompson, Mykhaylo Andriluka +2
cs.CVcs.LGcs.NEarXiv:1312.7302v62013Computed Tomography Reconstruction Using Deep Image Prior and Learned Reconstruction Methods
Daniel Otero Baguer, Johannes Leuschner, Maximilian Schmidt
eess.IVcs.CVcs.LGarXiv:2003.04989v22020Improving Position Encoding of Transformers for Multivariate Time Series Classification
Navid Mohammadi Foumani, Chang Wei Tan, Geoffrey I. Webb +1
cs.LGcs.CVarXiv:2305.16642v12023VCAR: Training-Free 3DGS Segmentation via View Completeness and Axis-Aware Boundary Refinement
Kun Cao, Di Wang, Haibin Zhu +4
cs.CVarXiv:2608.30870v12026Analytic Dynamics: Learning Physics-Grounded Representation for Fast Intrinsic Dynamics Inference from Monocular Videos
Jailing Lin, Jikuan Zhang, Jianhua Sun
cs.CVarXiv:2608.31025v12026Defense against Universal Adversarial Perturbations
Naveed Akhtar, Jian Liu, Ajmal Mian
cs.CVarXiv:1711.05929v32017Task-Driven Super Resolution: Object Detection in Low-resolution Images
Muhammad Haris, Greg Shakhnarovich, Norimichi Ukita
cs.CVarXiv:1803.11316v12018CDMamba: Incorporating Local Clues into Mamba for Remote Sensing Image Binary Change Detection
Haotian Zhang, Keyan Chen, Chenyang Liu +3
cs.CVarXiv:2406.04207v22024Modality Distillation with Multiple Stream Networks for Action Recognition
Nuno Garcia, Pietro Morerio, Vittorio Murino
cs.CVarXiv:1806.07110v22018Automatic calcium scoring in low-dose chest CT using deep neural networks with dilated convolutions
Nikolas Lessmann, Bram van Ginneken, Majd Zreik +4
cs.CVarXiv:1711.00349v22017Dense Optical Flow Prediction from a Static Image
Jacob Walker, Abhinav Gupta, Martial Hebert
cs.CVarXiv:1505.00295v22015Improving Referring Expression Grounding with Cross-modal Attention-guided Erasing
Xihui Liu, Zihao Wang, Jing Shao +2
cs.CVcs.CLarXiv:1903.00839v22019Music Gesture for Visual Sound Separation
Chuang Gan, Deng Huang, Hang Zhao +2
cs.CVcs.LGcs.MMarXiv:2004.09476v12020Learning Monocular Depth by Distilling Cross-domain Stereo Networks
Xiaoyang Guo, Hongsheng Li, Shuai Yi +2
cs.CVarXiv:1808.06586v12018Anatomy of Catastrophic Forgetting: Hidden Representations and Task Semantics
Vinay V. Ramasesh, Ethan Dyer, Maithra Raghu
cs.LGcs.CVstat.MLarXiv:2007.07400v12020Hierarchical Dynamic Filtering Network for RGB-D Salient Object Detection
Youwei Pang, Lihe Zhang, Xiaoqi Zhao +1
cs.CVarXiv:2007.06227v32020A Composition-Aware Pretraining Framework for Geospatial Foundation Models
Aryan Kashyap Naveen, Abhishek Srinivas, Pranav Moothedath +1
cs.CVcs.AIarXiv:2608.30817v12026Offboard 3D Object Detection from Point Cloud Sequences
Charles R. Qi, Yin Zhou, Mahyar Najibi +4
cs.CVarXiv:2103.05073v12021Evaluation of CNN-based Single-Image Depth Estimation Methods
Tobias Koch, Lukas Liebel, Friedrich Fraundorfer +1
cs.CVarXiv:1805.01328v12018Radiology Report Generation with a Learned Knowledge Base and Multi-modal Alignment
Shuxin Yang, Xian Wu, Shen Ge +2
eess.IVcs.CLcs.CVarXiv:2112.15011v22021Reading Car License Plates Using Deep Convolutional Neural Networks and LSTMs
Hui Li, Chunhua Shen
cs.CVarXiv:1601.05610v12016On Adversarial Robustness of Trajectory Prediction for Autonomous Vehicles
Qingzhao Zhang, Shengtuo Hu, Jiachen Sun +2
cs.CVcs.CRarXiv:2201.05057v32022FusionNet: 3D Object Classification Using Multiple Data Representations
Vishakh Hegde, Reza Zadeh
cs.CVarXiv:1607.05695v42016Dense Pose Transfer
Natalia Neverova, Riza Alp Guler, Iasonas Kokkinos
cs.CVarXiv:1809.01995v12018Real-Time Anomaly Detection and Localization in Crowded Scenes
Mohammad Sabokrou, Mahmood Fathy, Mojtaba Hosseini +1
cs.CVarXiv:1511.06936v12015Models Matter, So Does Training: An Empirical Study of CNNs for Optical Flow Estimation
Deqing Sun, Xiaodong Yang, Ming-Yu Liu +1
cs.CVarXiv:1809.05571v12018FANet: A Feedback Attention Network for Improved Biomedical Image Segmentation
Nikhil Kumar Tomar, Debesh Jha, Michael A. Riegler +5
cs.CVeess.IVarXiv:2103.17235v32021CORAL: A Benchmark for Structure-aware and Brain-wide Neuron Reconstruction in Light Microscopy
Zekang Yang, Jiamin Li, Zhenghua Li +3
cs.CVarXiv:2608.30768v12026Where are the Blobs: Counting by Localization with Point Supervision
Issam H. Laradji, Negar Rostamzadeh, Pedro O. Pinheiro +2
cs.CVarXiv:1807.09856v12018HorizonNet for visual terrain navigation
Bertil Grelsson, Andreas Robinson, Michael Felsberg +1
cs.CVcs.ROarXiv:2608.30471v12026Residual Attention U-Net for Automated Multi-Class Segmentation of COVID-19 Chest CT Images
Xiaocong Chen, Lina Yao, Yu Zhang
eess.IVcs.CVcs.LGarXiv:2004.05645v12020SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
Yang Zhan, Zhitong Xiong, Yuan Yuan
cs.CVarXiv:2401.09712v12024Motion-Saliency Complementary Masked Modeling for Point Cloud Video Understanding
Wei Wang, Yiding Sun, Yuyan Wang +4
cs.CVarXiv:2608.30279v12026Ground Truth Evaluation of Neural Network Explanations with CLEVR-XAI
Leila Arras, Ahmed Osman, Wojciech Samek
cs.CVcs.LGcs.NEarXiv:2003.07258v22020Visual Dexterity: In-Hand Reorientation of Novel and Complex Object Shapes
Tao Chen, Megha Tippur, Siyang Wu +3
cs.ROcs.AIcs.CVarXiv:2211.11744v32022Cube Padding for Weakly-Supervised Saliency Prediction in 360° Videos
Hsien-Tzu Cheng, Chun-Hung Chao, Jin-Dong Dong +3
cs.CVarXiv:1806.01320v12018Transcending Forgery Specificity with Latent Space Augmentation for Generalizable Deepfake Detection
Zhiyuan Yan, Yuhao Luo, Siwei Lyu +2
cs.CVarXiv:2311.11278v22023Multi-organ Segmentation over Partially Labeled Datasets with Multi-scale Feature Abstraction
Xi Fang, Pingkun Yan
cs.CVarXiv:2001.00208v22020PyMAF-X: Towards Well-aligned Full-body Model Regression from Monocular Images
Hongwen Zhang, Yating Tian, Yuxiang Zhang +4
cs.CVarXiv:2207.06400v42022