Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

11,941 to 12,000 of 18,817

  1. DPatch: An Adversarial Patch Attack on Object Detectors

    Xin Liu, Huanrui Yang, Ziwei Liu +3

    cs.CVcs.CRcs.LGarXiv:1806.02299v42018
  2. A Deep Cascade of Convolutional Neural Networks for MR Image Reconstruction

    Jo Schlemper, Jose Caballero, Joseph V. Hajnal +2

    cs.CVarXiv:1703.00555v12017
  3. SPINN: Synergistic Progressive Inference of Neural Networks over Device and Cloud

    Stefanos Laskaridis, Stylianos I. Venieris, Mario Almeida +2

    cs.LGcs.CVcs.DCarXiv:2008.06402v22020
  4. Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and Segmentation

    Jinyuan Liu, Zhu Liu, Guanyao Wu +5

    cs.CVarXiv:2308.02097v12023
  5. Multi-level Feature Learning for Contrastive Multi-view Clustering

    Jie Xu, Huayi Tang, Yazhou Ren +3

    cs.LGcs.CVarXiv:2106.11193v22021
  6. Online Multi-Object Tracking Using CNN-based Single Object Tracker with Spatial-Temporal Attention Mechanism

    Qi Chu, Wanli Ouyang, Hongsheng Li +3

    cs.CVarXiv:1708.02843v22017
  7. MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

    Bin Lin, Zhenyu Tang, Yang Ye +7

    cs.CVarXiv:2401.15947v52024
  8. OpenCFU, a New Free and Open-Source Software to Count Cell Colonies and Other Circular Objects

    Quentin Geissmann

    q-bio.QMcs.CVarXiv:1210.5502v32012
  9. Real-Time Intermediate Flow Estimation for Video Frame Interpolation

    Zhewei Huang, Tianyuan Zhang, Wen Heng +2

    cs.CVcs.LGarXiv:2011.06294v122020
  10. Imbalanced Deep Learning by Minority Class Incremental Rectification

    Qi Dong, Shaogang Gong, Xiatian Zhu

    cs.CVarXiv:1804.10851v12018
  11. Regularization for Deep Learning: A Taxonomy

    Jan Kukačka, Vladimir Golkov, Daniel Cremers

    cs.LGcs.AIcs.CVarXiv:1710.10686v12017
  12. Inverting The Generator Of A Generative Adversarial Network

    Antonia Creswell, Anil Anthony Bharath

    cs.CVcs.LGarXiv:1611.05644v12016
  13. Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving

    Long Chen, Oleg Sinavski, Jan Hünermann +5

    cs.ROcs.AIcs.CLarXiv:2310.01957v22023
  14. InceptionNeXt: When Inception Meets ConvNeXt

    Weihao Yu, Pan Zhou, Shuicheng Yan +1

    cs.CVcs.AIcs.LGarXiv:2303.16900v32023
  15. Cross-Modality Paired-Images Generation for RGB-Infrared Person Re-Identification

    Guan-An Wang, Tianzhu Zhang. Yang Yang, Jian Cheng +3

    cs.CVarXiv:2002.04114v22020
  16. DetCo: Unsupervised Contrastive Learning for Object Detection

    Enze Xie, Jian Ding, Wenhai Wang +5

    cs.CVarXiv:2102.04803v22021
  17. UACANet: Uncertainty Augmented Context Attention for Polyp Segmentation

    Taehun Kim, Hyemin Lee, Daijin Kim

    cs.CVarXiv:2107.02368v32021
  18. Extreme Parkour with Legged Robots

    Xuxin Cheng, Kexin Shi, Ananye Agarwal +1

    cs.ROcs.AIcs.CVarXiv:2309.14341v12023
  19. Multimodal Virtual Point 3D Detection

    Tianwei Yin, Xingyi Zhou, Philipp Krähenbühl

    cs.CVcs.LGcs.ROarXiv:2111.06881v12021
  20. AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection

    Thanh-Toan Do, Anh Nguyen, Ian Reid

    cs.CVcs.ROarXiv:1709.07326v32017
  21. DIODE: A Dense Indoor and Outdoor DEpth Dataset

    Igor Vasiljevic, Nick Kolkin, Shanyi Zhang +8

    cs.CVarXiv:1908.00463v22019
  22. Learning Image-adaptive 3D Lookup Tables for High Performance Photo Enhancement in Real-time

    Hui Zeng, Jianrui Cai, Lida Li +2

    eess.IVcs.CVarXiv:2009.14468v12020
  23. UniT: Multimodal Multitask Learning with a Unified Transformer

    Ronghang Hu, Amanpreet Singh

    cs.CVcs.CLarXiv:2102.10772v32021
  24. Skip Connections Matter: On the Transferability of Adversarial Examples Generated with ResNets

    Dongxian Wu, Yisen Wang, Shu-Tao Xia +2

    cs.LGcs.CRcs.CVarXiv:2002.05990v12020
  25. Sparse Inertial Poser: Automatic 3D Human Pose Estimation from Sparse IMUs

    Timo von Marcard, Bodo Rosenhahn, Michael J. Black +1

    cs.CVcs.GRarXiv:1703.08014v22017
  26. RTM3D: Real-time Monocular 3D Detection from Object Keypoints for Autonomous Driving

    Peixuan Li, Huaici Zhao, Pengfei Liu +1

    cs.CVcs.ROeess.IVarXiv:2001.03343v12020
  27. Online Multi-Object Tracking with Dual Matching Attention Networks

    Ji Zhu, Hua Yang, Nian Liu +3

    cs.CVarXiv:1902.00749v12019
  28. In Defense of Grid Features for Visual Question Answering

    Huaizu Jiang, Ishan Misra, Marcus Rohrbach +2

    cs.CVarXiv:2001.03615v22020
  29. 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

    Tsung-Wei Ke, Nikolaos Gkanatsios, Katerina Fragkiadaki

    cs.ROcs.AIcs.CVarXiv:2402.10885v32024
  30. Visual Prompt Multi-Modal Tracking

    Jiawen Zhu, Simiao Lai, Xin Chen +2

    cs.CVarXiv:2303.10826v22023
  31. Rethinking Semantic Segmentation: A Prototype View

    Tianfei Zhou, Wenguan Wang, Ender Konukoglu +1

    cs.CVarXiv:2203.15102v22022
  32. Swapping Autoencoder for Deep Image Manipulation

    Taesung Park, Jun-Yan Zhu, Oliver Wang +4

    cs.CVcs.GRcs.LGarXiv:2007.00653v22020
  33. NoPe-NeRF: Optimising Neural Radiance Field with No Pose Prior

    Wenjing Bian, Zirui Wang, Kejie Li +2

    cs.CVarXiv:2212.07388v32022
  34. Full-Gradient Representation for Neural Network Visualization

    Suraj Srinivas, Francois Fleuret

    cs.LGcs.CVstat.MLarXiv:1905.00780v42019
  35. Deep filter banks for texture recognition, description, and segmentation

    Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos +1

    cs.CVarXiv:1507.02620v22015
  36. Kimera: from SLAM to Spatial Perception with 3D Dynamic Scene Graphs

    Antoni Rosinol, Andrew Violette, Marcus Abate +5

    cs.ROcs.CVarXiv:2101.06894v32021
  37. AutoTrack: Towards High-Performance Visual Tracking for UAV with Automatic Spatio-Temporal Regularization

    Yiming Li, Changhong Fu, Fangqiang Ding +2

    cs.CVarXiv:2003.12949v12020
  38. The Unreasonable Effectiveness of Noisy Data for Fine-Grained Recognition

    Jonathan Krause, Benjamin Sapp, Andrew Howard +5

    cs.CVarXiv:1511.06789v32015
  39. STGAN: A Unified Selective Transfer Network for Arbitrary Image Attribute Editing

    Ming Liu, Yukang Ding, Min Xia +4

    cs.CVarXiv:1904.09709v12019
  40. Emotions Don't Lie: An Audio-Visual Deepfake Detection Method Using Affective Cues

    Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra +2

    cs.CVcs.LGcs.SDarXiv:2003.06711v32020
  41. An Ensemble-based System for Microaneurysm Detection and Diabetic Retinopathy Grading

    Balint Antal, Andras Hajdu

    cs.CVcs.AIstat.AParXiv:1410.8577v12014
  42. Triplet-Center Loss for Multi-View 3D Object Retrieval

    Xinwei He, Yang Zhou, Zhichao Zhou +2

    cs.CVarXiv:1803.06189v12018
  43. Image Segmentation by Using Threshold Techniques

    Salem Saleh Al-amri, N. V. Kalyankar, Khamitkar S. D.

    cs.CVarXiv:1005.4020v12010
    Summaries:한국어
  44. Convolutional Sequence to Sequence Model for Human Dynamics

    Chen Li, Zhen Zhang, Wee Sun Lee +1

    cs.CVarXiv:1805.00655v12018
  45. TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval

    Jie Lei, Licheng Yu, Tamara L. Berg +1

    cs.CVcs.CLcs.IRarXiv:2001.09099v22020
  46. Improving Sign Language Translation with Monolingual Data by Sign Back-Translation

    Hao Zhou, Wengang Zhou, Weizhen Qi +2

    cs.CVcs.CLarXiv:2105.12397v12021
  47. The Devil is in the Channels: Mutual-Channel Loss for Fine-Grained Image Classification

    Dongliang Chang, Yifeng Ding, Jiyang Xie +6

    cs.CVarXiv:2002.04264v32020
  48. Learning to Learn from Noisy Labeled Data

    Junnan Li, Yongkang Wong, Qi Zhao +1

    cs.LGcs.CVstat.MLarXiv:1812.05214v22018
  49. Common Diffusion Noise Schedules and Sample Steps are Flawed

    Shanchuan Lin, Bingchen Liu, Jiashi Li +1

    cs.CVarXiv:2305.08891v42023
  50. CANet: Cross-disease Attention Network for Joint Diabetic Retinopathy and Diabetic Macular Edema Grading

    Xiaomeng Li, Xiaowei Hu, Lequan Yu +3

    eess.IVcs.CVarXiv:1911.01376v12019
  51. Transforming Model Prediction for Tracking

    Christoph Mayer, Martin Danelljan, Goutam Bhat +4

    cs.CVarXiv:2203.11192v12022
  52. Student-Teacher Feature Pyramid Matching for Anomaly Detection

    Guodong Wang, Shumin Han, Errui Ding +1

    cs.CVarXiv:2103.04257v32021
  53. Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-training

    Weituo Hao, Chunyuan Li, Xiujun Li +2

    cs.CVcs.CLcs.LGarXiv:2002.10638v22020
  54. Image Classification with Deep Learning in the Presence of Noisy Labels: A Survey

    Görkem Algan, Ilkay Ulusoy

    cs.LGcs.CVstat.MLarXiv:1912.05170v32019
  55. Foreground-aware Image Inpainting

    Wei Xiong, Jiahui Yu, Zhe Lin +4

    cs.CVarXiv:1901.05945v32019
  56. SEN12MS -- A Curated Dataset of Georeferenced Multi-Spectral Sentinel-1/2 Imagery for Deep Learning and Data Fusion

    Michael Schmitt, Lloyd Haydn Hughes, Chunping Qiu +1

    cs.CVarXiv:1906.07789v12019
  57. Multi-Scale Spatial Temporal Graph Convolutional Network for Skeleton-Based Action Recognition

    Zhan Chen, Sicheng Li, Bing Yang +2

    cs.CVarXiv:2206.13028v12022
  58. Yin and Yang: Balancing and Answering Binary Visual Questions

    Peng Zhang, Yash Goyal, Douglas Summers-Stay +2

    cs.CLcs.CVcs.LGarXiv:1511.05099v52015
  59. EvalCrafter: Benchmarking and Evaluating Large Video Generation Models

    Yaofang Liu, Xiaodong Cun, Xuebo Liu +7

    cs.CVarXiv:2310.11440v32023
  60. OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

    Hugo Laurençon, Lucile Saulnier, Léo Tronchon +9

    cs.IRcs.CVarXiv:2306.16527v22023