Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
11,281 to 11,340 of 18,839
Clean-Label Backdoor Attacks on Video Recognition Models
Shihao Zhao, Xingjun Ma, Xiang Zheng +3
cs.CVarXiv:2003.03030v22020Visual Dialog
Abhishek Das, Satwik Kottur, Khushi Gupta +5
cs.CVcs.AIcs.CLarXiv:1611.08669v52016Fully Convolutional Instance-aware Semantic Segmentation
Yi Li, Haozhi Qi, Jifeng Dai +2
cs.CVarXiv:1611.07709v22016DexPilot: Vision Based Teleoperation of Dexterous Robotic Hand-Arm System
Ankur Handa, Karl Van Wyk, Wei Yang +6
cs.CVcs.LGcs.ROarXiv:1910.03135v22019Driving in the Matrix: Can Virtual Worlds Replace Human-Generated Annotations for Real World Tasks?
Matthew Johnson-Roberson, Charles Barto, Rounak Mehta +3
cs.CVcs.ROarXiv:1610.01983v22016Full Resolution Image Compression with Recurrent Neural Networks
George Toderici, Damien Vincent, Nick Johnston +4
cs.CVarXiv:1608.05148v22016Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising
Kai Zhang, Wangmeng Zuo, Yunjin Chen +2
cs.CVarXiv:1608.03981v12016R-FCN: Object Detection via Region-based Fully Convolutional Networks
Jifeng Dai, Yi Li, Kaiming He +1
cs.CVarXiv:1605.06409v32016A Haar Wavelet-Based Perceptual Similarity Index for Image Quality Assessment
Rafael Reisenhofer, Sebastian Bosse, Gitta Kutyniok +1
cs.CVarXiv:1607.06140v42016FD-GAN: Generative Adversarial Networks with Fusion-discriminator for Single Image Dehazing
Yu Dong, Yihao Liu, He Zhang +2
cs.CVarXiv:2001.06968v220203DMatch: Learning Local Geometric Descriptors from RGB-D Reconstructions
Andy Zeng, Shuran Song, Matthias Nießner +3
cs.CVarXiv:1603.08182v32016Masked Visual Pre-training for Motor Control
Tete Xiao, Ilija Radosavovic, Trevor Darrell +1
cs.CVcs.LGcs.ROarXiv:2203.06173v12022Combining Markov Random Fields and Convolutional Neural Networks for Image Synthesis
Chuan Li, Michael Wand
cs.CVarXiv:1601.04589v12016Towards Open Set Deep Networks
Abhijit Bendale, Terrance Boult
cs.CVcs.LGarXiv:1511.06233v12015STC: A Simple to Complex Framework for Weakly-supervised Semantic Segmentation
Yunchao Wei, Xiaodan Liang, Yunpeng Chen +5
cs.CVarXiv:1509.03150v22015Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross Girshick +1
cs.CVarXiv:1506.01497v32015Learning to rank in person re-identification with metric ensembles
Sakrapee Paisitkriangkrai, Chunhua Shen, Anton van den Hengel
cs.CVarXiv:1503.01543v12015BEHAVE: Dataset and Method for Tracking Human Object Interactions
Bharat Lal Bhatnagar, Xianghui Xie, Ilya A. Petrov +3
cs.CVarXiv:2204.06950v12022Computing the Stereo Matching Cost with a Convolutional Neural Network
Jure Žbontar, Yann LeCun
cs.CVcs.LGcs.NEarXiv:1409.4326v22014Learn Convolutional Neural Network for Face Anti-Spoofing
Jianwei Yang, Zhen Lei, Stan Z. Li
cs.CVarXiv:1408.5601v22014FSS-1000: A 1000-Class Dataset for Few-Shot Segmentation
Xiang Li, Tianhan Wei, Yau Pun Chen +2
cs.CVarXiv:1907.12347v22019Scalable Object Detection using Deep Neural Networks
Dumitru Erhan, Christian Szegedy, Alexander Toshev +1
cs.CVstat.MLarXiv:1312.2249v12013Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell +1
cs.CVarXiv:1311.2524v52013Pedestrian Detection with Unsupervised Multi-Stage Feature Learning
Pierre Sermanet, Koray Kavukcuoglu, Soumith Chintala +1
cs.CVcs.LGarXiv:1212.0142v22012PhysFormer: Facial Video-based Physiological Measurement with Temporal Difference Transformer
Zitong Yu, Yuming Shen, Jingang Shi +3
cs.CVarXiv:2111.12082v22021A Transformer-based representation-learning model with unified processing of multimodal input for clinical diagnostics
Hong-Yu Zhou, Yizhou Yu, Chengdi Wang +7
cs.CVcs.CLcs.LGarXiv:2306.00864v12023A Comprehensive Survey on Pose-Invariant Face Recognition
Changxing Ding, Dacheng Tao
cs.CVarXiv:1502.04383v32015Automatic Detection of Coronavirus Disease (COVID-19) in X-ray and CT Images: A Machine Learning-Based Approach
Sara Hosseinzadeh Kassani, Peyman Hosseinzadeh Kassasni, Michal J. Wesolowski +2
eess.IVcs.CVarXiv:2004.10641v12020Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Junyang Wang, Haiyang Xu, Jiabo Ye +5
cs.CLcs.CVarXiv:2401.16158v22024DETRs with Hybrid Matching
Ding Jia, Yuhui Yuan, Haodi He +6
cs.CVarXiv:2207.13080v32022Geography-Aware Self-Supervised Learning
Kumar Ayush, Burak Uzkent, Chenlin Meng +4
cs.CVarXiv:2011.09980v72020Backbone is All Your Need: A Simplified Architecture for Visual Object Tracking
Boyu Chen, Peixia Li, Lei Bai +6
cs.CVarXiv:2203.05328v22022VIGOR: Cross-View Image Geo-localization beyond One-to-one Retrieval
Sijie Zhu, Taojiannan Yang, Chen Chen
cs.CVcs.AIarXiv:2011.12172v22020Generalized Category Discovery
Sagar Vaze, Kai Han, Andrea Vedaldi +1
cs.CVcs.LGarXiv:2201.02609v22022FreeAnchor: Learning to Match Anchors for Visual Object Detection
Xiaosong Zhang, Fang Wan, Chang Liu +2
cs.CVcs.LGarXiv:1909.02466v22019Emergence of Language with Multi-agent Games: Learning to Communicate with Sequences of Symbols
Serhii Havrylov, Ivan Titov
cs.LGcs.CLcs.CVarXiv:1705.11192v22017Adaptive Context Selection for Polyp Segmentation
Ruifei Zhang, Guanbin Li, Zhen Li +3
cs.CVarXiv:2301.04799v12023Unsupervised Domain Adaptation via Structurally Regularized Deep Clustering
Hui Tang, Ke Chen, Kui Jia
cs.CVarXiv:2003.08607v12020COTR: Correspondence Transformer for Matching Across Images
Wei Jiang, Eduard Trulls, Jan Hosang +2
cs.CVarXiv:2103.14167v22021RoboTHOR: An Open Simulation-to-Real Embodied AI Platform
Matt Deitke, Winson Han, Alvaro Herrasti +10
cs.CVcs.ROarXiv:2004.06799v12020Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models
Songwei Ge, Seungjun Nah, Guilin Liu +7
cs.CVcs.GRcs.LGarXiv:2305.10474v32023STAR: A Benchmark for Situated Reasoning in Real-World Videos
Bo Wu, Shoubin Yu, Zhenfang Chen +2
cs.AIcs.CLcs.CVarXiv:2405.09711v12024Hierarchical Bilinear Pooling for Fine-Grained Visual Recognition
Chaojian Yu, Xinyi Zhao, Qi Zheng +2
cs.CVarXiv:1807.09915v12018Parametric Noise Injection: Trainable Randomness to Improve Deep Neural Network Robustness against Adversarial Attack
Adnan Siraj Rakin, Zhezhi He, Deliang Fan
cs.LGcs.CRcs.CVarXiv:1811.09310v12018Toward Fast, Flexible, and Robust Low-Light Image Enhancement
Long Ma, Tengyu Ma, Risheng Liu +2
cs.CVarXiv:2204.10137v12022Bringing a Blurry Frame Alive at High Frame-Rate with an Event Camera
Liyuan Pan, Cedric Scheerlinck, Xin Yu +3
cs.CVarXiv:1811.10180v22018Deep Unsupervised Domain Adaptation: A Review of Recent Advances and Perspectives
Xiaofeng Liu, Chaehwa Yoo, Fangxu Xing +4
cs.CVcs.AIcs.LGarXiv:2208.07422v12022Improved YOLOv5 network for real-time multi-scale traffic sign detection
Junfan Wang, Yi Chen, Mingyu Gao +1
cs.CVcs.LGarXiv:2112.08782v22021Structured Graph Learning for Scalable Subspace Clustering: From Single-view to Multi-view
Zhao Kang, Zhiping Lin, Xiaofeng Zhu +1
cs.LGcs.AIcs.CVarXiv:2102.07943v12021Deep Feature Interpolation for Image Content Changes
Paul Upchurch, Jacob Gardner, Geoff Pleiss +4
cs.CVarXiv:1611.05507v22016Pixel-Level Domain Transfer
Donggeun Yoo, Namil Kim, Sunggyun Park +2
cs.CVcs.AIarXiv:1603.07442v32016DeepICP: An End-to-End Deep Neural Network for 3D Point Cloud Registration
Weixin Lu, Guowei Wan, Yao Zhou +3
cs.CVcs.CGcs.GRarXiv:1905.04153v22019Zero-Shot Visual Imitation
Deepak Pathak, Parsa Mahmoudieh, Guanghao Luo +7
cs.LGcs.AIcs.CVarXiv:1804.08606v12018Changer: Feature Interaction is What You Need for Change Detection
Sheng Fang, Kaiyu Li, Zhe Li
cs.CVarXiv:2209.08290v12022TrafficSim: Learning to Simulate Realistic Multi-Agent Behaviors
Simon Suo, Sebastian Regalado, Sergio Casas +1
cs.ROcs.AIcs.CVarXiv:2101.06557v12021Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with Transformers
Zi-Xin Zou, Zhipeng Yu, Yuan-Chen Guo +4
cs.CVarXiv:2312.09147v22023On-Device Training Under 256KB Memory
Ji Lin, Ligeng Zhu, Wei-Ming Chen +3
cs.CVarXiv:2206.15472v42022MACE: Mass Concept Erasure in Diffusion Models
Shilin Lu, Zilan Wang, Leyang Li +2
cs.CVcs.AIcs.LGarXiv:2403.06135v12024Deep Learning and Machine Vision for Food Processing: A Survey
Lili Zhu, Petros Spachos, Erica Pensini +1
cs.CVcs.LGarXiv:2103.16106v12021Rotate your Networks: Better Weight Consolidation and Less Catastrophic Forgetting
Xialei Liu, Marc Masana, Luis Herranz +3
cs.CVarXiv:1802.02950v42018