Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

661 to 720 of 18,845

  1. Celeb-DF: A Large-scale Challenging Dataset for DeepFake Forensics

    Yuezun Li, Xin Yang, Pu Sun +2

    cs.CRcs.CVeess.IVarXiv:1909.12962v42019
  2. Synthetic Data for Deep Learning

    Sergey I. Nikolenko

    cs.LGcs.CRcs.CVarXiv:1909.11512v12019
  3. DE-FAKE: Detection and Attribution of Fake Images Generated by Text-to-Image Generation Models

    Zeyang Sha, Zheng Li, Ning Yu +1

    cs.CRcs.CVcs.LGarXiv:2210.06998v22022
  4. From Points to Parts: 3D Object Detection from Point Cloud with Part-aware and Part-aggregation Network

    Shaoshuai Shi, Zhe Wang, Jianping Shi +2

    cs.CVarXiv:1907.03670v32019
  5. Physical Adversarial Attack meets Computer Vision: A Decade Survey

    Hui Wei, Hao Tang, Xuemei Jia +6

    cs.CVarXiv:2209.15179v42022
  6. On Physical Adversarial Patches for Object Detection

    Mark Lee, Zico Kolter

    cs.CVcs.CRcs.LGarXiv:1906.11897v12019
  7. LUX: A Lesion-Aware Graph-Conditioned Visual - Language Architecture for Explainable Endoscopic Captioning

    Alexis Ivan Escamilla-Lopez, Gilberto Ochoa-Ruiz, Salvador Hinojosa +1

    cs.CVarXiv:2608.23853v12026
  8. Automatic Colon Polyp Detection using Region based Deep CNN and Post Learning Approaches

    Younghak Shin, Hemin Ali Qadir, Lars Aabakken +2

    cs.CVcs.AIarXiv:1906.11463v12019
  9. Video In Sentences Out

    Andrei Barbu, Alexander Bridge, Zachary Burchill +15

    cs.CVcs.CLcs.IRarXiv:1408.6418v12014
  10. NAS-FCOS: Fast Neural Architecture Search for Object Detection

    Ning Wang, Yang Gao, Hao Chen +4

    cs.CVarXiv:1906.04423v42019
  11. Analyzing the Performance of Multilayer Neural Networks for Object Recognition

    Pulkit Agrawal, Ross Girshick, Jitendra Malik

    cs.CVcs.NEarXiv:1407.1610v22014
  12. Rethinking the Unpretentious U-net for Medical Ultrasound Image Segmentation

    Gongping Chen, Lei Li, JianXun Zhang +1

    eess.IVcs.CVcs.LGarXiv:2209.07193v42022
  13. YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications

    Chuyi Li, Lulu Li, Hongliang Jiang +15

    cs.CVarXiv:2209.02976v12022
  14. Rigid-Motion Scattering for Texture Classification

    Laurent SIfre, Stéphane Mallat

    cs.CVarXiv:1403.1687v12014
  15. Deep Patch Visual Odometry

    Zachary Teed, Lahav Lipson, Jia Deng

    cs.CVarXiv:2208.04726v22022
  16. Improving Robustness Without Sacrificing Accuracy with Patch Gaussian Augmentation

    Raphael Gontijo Lopes, Dong Yin, Ben Poole +2

    cs.LGcs.CVstat.MLarXiv:1906.02611v12019
  17. Zero-Shot Semantic Segmentation

    Maxime Bucher, Tuan-Hung Vu, Matthieu Cord +1

    cs.CVarXiv:1906.00817v22019
  18. VeCAS: Vessel-Focused Contrast-Free Angiogram Synthesis for Vascular Interventions

    De-Xing Huang, Chen-Yu Wang, Hao Liang +10

    cs.CVarXiv:2608.22828v12026
  19. Are Disentangled Representations Helpful for Abstract Visual Reasoning?

    Sjoerd van Steenkiste, Francesco Locatello, Jürgen Schmidhuber +1

    cs.LGcs.CVcs.NEarXiv:1905.12506v32019
  20. Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

    Jiahui Yu, Yuanzhong Xu, Jing Yu Koh +14

    cs.CVcs.LGarXiv:2206.10789v12022
  21. Multiview Hessian Discriminative Sparse Coding for Image Annotation

    Weifeng Liu, Dacheng Tao, Jun Cheng +1

    cs.MMcs.CVcs.ITarXiv:1307.3811v12013
  22. DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation

    Fenghao Lei, Zhixiong Huang, Long Yang +5

    cs.ROcs.AIcs.CVarXiv:2608.22067v32026
  23. Zero-Shot Knowledge Distillation in Deep Networks

    Gaurav Kumar Nayak, Konda Reddy Mopuri, Vaisakh Shaj +2

    cs.LGcs.CVstat.MLarXiv:1905.08114v12019
  24. Slim-neck by GSConv: A lightweight-design for real-time detector architectures

    Hulin Li, Jun Li, Hanbing Wei +3

    cs.CVarXiv:2206.02424v32022
  25. Recurrent Video Restoration Transformer with Guided Deformable Attention

    Jingyun Liang, Yuchen Fan, Xiaoyu Xiang +7

    cs.CVeess.IVarXiv:2206.02146v32022
  26. A Model or 603 Exemplars: Towards Memory-Efficient Class-Incremental Learning

    Da-Wei Zhou, Qi-Wei Wang, Han-Jia Ye +1

    cs.LGcs.CVarXiv:2205.13218v22022
  27. Decoupled-and-Coupled Networks: Self-Supervised Hyperspectral Image Super-Resolution with Subpixel Fusion

    Danfeng Hong, Jing Yao, Deyu Meng +2

    eess.IVcs.CVarXiv:2205.03742v12022
  28. Generating Multiple Hypotheses for 3D Human Pose Estimation with Mixture Density Network

    Chen Li, Gim Hee Lee

    cs.CVarXiv:1904.05547v12019
  29. Depth from Videos in the Wild: Unsupervised Monocular Depth Learning from Unknown Cameras

    Ariel Gordon, Hanhan Li, Rico Jonschkowski +1

    cs.CVcs.GRcs.LGarXiv:1904.04998v12019
  30. Visible-Thermal UAV Tracking: A Large-Scale Benchmark and New Baseline

    Pengyu Zhang, Jie Zhao, Dong Wang +2

    cs.CVarXiv:2204.04120v12022
  31. Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion

    Bowen Cui, Weijie Wang, Zeyu Zhang +7

    cs.CVarXiv:2608.19567v32026
  32. UG$^{2+}$ Track 2: A Collective Benchmark Effort for Evaluating and Advancing Image Understanding in Poor Visibility Environments

    Ye Yuan, Wenhan Yang, Wenqi Ren +3

    cs.CVarXiv:1904.04474v42019
  33. Towards An End-to-End Framework for Flow-Guided Video Inpainting

    Zhen Li, Cheng-Ze Lu, Jianhua Qin +2

    eess.IVcs.CVarXiv:2204.02663v22022
  34. ConvPoint: Continuous Convolutions for Point Cloud Processing

    Alexandre Boulch

    cs.CVarXiv:1904.02375v52019
  35. AdaFace: Quality Adaptive Margin for Face Recognition

    Minchul Kim, Anil K. Jain, Xiaoming Liu

    cs.CVarXiv:2204.00964v22022
  36. TensorMask: A Foundation for Dense Object Segmentation

    Xinlei Chen, Ross Girshick, Kaiming He +1

    cs.CVarXiv:1903.12174v22019
  37. Photorealistic Style Transfer via Wavelet Transforms

    Jaejun Yoo, Youngjung Uh, Sanghyuk Chun +2

    cs.CVarXiv:1903.09760v22019
  38. Constrained Few-shot Class-incremental Learning

    Michael Hersche, Geethan Karunaratne, Giovanni Cherubini +3

    cs.CVcs.LGarXiv:2203.16588v12022
  39. OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object Interaction

    Lixin Yang, Kailin Li, Xinyu Zhan +4

    cs.CVarXiv:2203.15709v12022
  40. Deconvolving Images with Unknown Boundaries Using the Alternating Direction Method of Multipliers

    Mariana S. C. Almeida, Mário A. T. Figueiredo

    math.OCcs.CVarXiv:1210.2687v22012
  41. Predictive Inequity in Object Detection

    Benjamin Wilson, Judy Hoffman, Jamie Morgenstern

    cs.CVcs.LGstat.MLarXiv:1902.11097v12019
  42. Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework

    Botao Ye, Hong Chang, Bingpeng Ma +2

    cs.CVarXiv:2203.11991v42022
  43. Improving Generalization in Federated Learning by Seeking Flat Minima

    Debora Caldarola, Barbara Caputo, Marco Ciccone

    cs.LGcs.CVarXiv:2203.11834v32022
  44. Leveraging existing sparse point annotations for benthic imagery dense segmentation

    Cesar Borja, Breck A. McCollum, Jarret E. Byrnes +2

    cs.CVcs.LGarXiv:2608.17561v12026
  45. Skeleton-Based Online Action Prediction Using Scale Selection Network

    Jun Liu, Amir Shahroudy, Gang Wang +2

    cs.CVarXiv:1902.03084v22019
  46. PixRestore: Unified Image Restoration via Pixel Diffusion Transformer

    Lingchen Sun, Rongyuan Wu, Xiangtao Kong +6

    cs.CVarXiv:2608.16793v12026
  47. Implicit 3D Orientation Learning for 6D Object Detection from RGB Images

    Martin Sundermeyer, Zoltan-Csaba Marton, Maximilian Durner +2

    cs.CVarXiv:1902.01275v22019
  48. ABAW: Valence-Arousal Estimation, Expression Recognition, Action Unit Detection & Multi-Task Learning Challenges

    Dimitrios Kollias

    cs.CVcs.LGarXiv:2202.10659v22022
  49. Sparse Subspace Clustering: Algorithm, Theory, and Applications

    Ehsan Elhamifar, Rene Vidal

    cs.CVcs.IRcs.ITarXiv:1203.1005v32012
  50. Patches Are All You Need?

    Asher Trockman, J. Zico Kolter

    cs.CVcs.AIcs.LGarXiv:2201.09792v12022
  51. Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents

    Wenlong Huang, Pieter Abbeel, Deepak Pathak +1

    cs.LGcs.AIcs.CLarXiv:2201.07207v22022
  52. DeCo-MIL: Debiased Counterfactual Reasoning for Long-Tailed Whole Slide Image Analysis

    Xiaoxiao Li, Xitong Ling, Jiawen Li +6

    cs.CVcs.AIarXiv:2608.14719v12026
  53. Compressive Imaging using Approximate Message Passing and a Markov-Tree Prior

    Subhojit Som, Philip Schniter

    cs.CVarXiv:1108.2632v12011
  54. NeRF for Outdoor Scene Relighting

    Viktor Rudnev, Mohamed Elgharib, William Smith +3

    cs.CVcs.GRarXiv:2112.05140v22021
  55. Inverse Cooking: Recipe Generation from Food Images

    Amaia Salvador, Michal Drozdzal, Xavier Giro-i-Nieto +1

    cs.CVarXiv:1812.06164v22018
  56. Prompting Visual-Language Models for Efficient Video Understanding

    Chen Ju, Tengda Han, Kunhao Zheng +2

    cs.CVcs.CLarXiv:2112.04478v22021
  57. A fast accurate fine-grain object detection model based on YOLOv4 deep neural network

    Arunabha M. Roy, Rikhi Bose, Jayabrata Bhaduri

    cs.CVcs.LGarXiv:2111.00298v12021
  58. Improving Robustness using Generated Data

    Sven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles +3

    cs.LGcs.CVstat.MLarXiv:2110.09468v22021
  59. Self-supervised Learning is More Robust to Dataset Imbalance

    Hong Liu, Jeff Z. HaoChen, Adrien Gaidon +1

    cs.LGcs.CVstat.MLarXiv:2110.05025v22021
  60. Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

    Zihan Wang, Seungjun Lee, Yinghao Xu +1

    cs.CVcs.ROarXiv:2607.05765v12026