Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,921 to 1,980 of 18,855

  1. A DenseNet Based Approach for Multi-Frame In-Loop Filter in HEVC

    Tianyi Li, Mai Xu, Ren Yang +1

    cs.CVarXiv:1903.01648v12019
  2. Vision-Language Models as Success Detectors

    Yuqing Du, Ksenia Konyushkova, Misha Denil +5

    cs.CVcs.AIcs.LGarXiv:2303.07280v12023
  3. Application of Deep Learning in Fundus Image Processing for Ophthalmic Diagnosis -- A Review

    Sourya Sengupta, Amitojdeep Singh, Henry A. Leopold +2

    cs.CVcs.LGstat.MLarXiv:1812.07101v32018
  4. GenSim: Generating Robotic Simulation Tasks via Large Language Models

    Lirui Wang, Yiyang Ling, Zhecheng Yuan +6

    cs.LGcs.CLcs.CVarXiv:2310.01361v22023
  5. End-to-end Active Object Tracking and Its Real-world Deployment via Reinforcement Learning

    Wenhan Luo, Peng Sun, Fangwei Zhong +3

    cs.CVarXiv:1808.03405v22018
  6. CRAFT: Camera-Radar 3D Object Detection with Spatio-Contextual Fusion Transformer

    Youngseok Kim, Sanmin Kim, Jun Won Choi +1

    cs.CVcs.AIcs.ROarXiv:2209.06535v22022
  7. Agile But Safe: Learning Collision-Free High-Speed Legged Locomotion

    Tairan He, Chong Zhang, Wenli Xiao +3

    cs.ROcs.AIcs.CVarXiv:2401.17583v32024
  8. CLIP-Event: Connecting Text and Images with Event Structures

    Manling Li, Ruochen Xu, Shuohang Wang +6

    cs.CVcs.AIarXiv:2201.05078v22022
  9. Fast Unsupervised Brain Anomaly Detection and Segmentation with Diffusion Models

    Walter H. L. Pinaya, Mark S. Graham, Robert Gray +12

    cs.CVeess.IVq-bio.QMarXiv:2206.03461v12022
  10. Geodesic Exponential Kernels: When Curvature and Linearity Conflict

    Aasa Feragen, Francois Lauze, Søren Hauberg

    cs.LGcs.CVarXiv:1411.0296v22014
  11. ISIA Food-500: A Dataset for Large-Scale Food Recognition via Stacked Global-Local Attention Network

    Weiqing Min, Linhu Liu, Zhiling Wang +4

    cs.CVcs.MMarXiv:2008.05655v12020
  12. SA-Profile: Automated Sulcus Angle Profiling from Super-Resolution MRI

    Michael Wehrli, Leo Widmer, Edwin Li +5

    cs.CVcs.AIarXiv:2609.10125v12026
  13. Local Class-Specific and Global Image-Level Generative Adversarial Networks for Semantic-Guided Scene Generation

    Hao Tang, Dan Xu, Yan Yan +2

    cs.CVcs.LGeess.IVarXiv:1912.12215v32019
  14. A Comprehensive Review of Data-Driven Co-Speech Gesture Generation

    Simbarashe Nyatsanga, Taras Kucherenko, Chaitanya Ahuja +2

    cs.GRcs.CVcs.HCarXiv:2301.05339v42023
  15. ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe

    Yifan Bai, Zeyang Zhao, Yihong Gong +1

    cs.CVarXiv:2312.17133v32023
  16. A statistical approach to bias in zero-shot learning: the lens of handwriting recognition

    Clarence Chew, Gim Siang Chia, Sukalpa Chanda +2

    stat.MLcs.AIcs.CVarXiv:2609.10084v12026
  17. The Multi-modality Cell Segmentation Challenge: Towards Universal Solutions

    Jun Ma, Ronald Xie, Shamini Ayyadhury +37

    eess.IVcs.CVcs.LGarXiv:2308.05864v22023
  18. Elastoformer: Enabling Dynamic Adaptivity via Elastic Model Transformation

    Sudaksh Kalra, Dolly Sapra

    cs.CVcs.AIcs.PFarXiv:2609.10018v12026
  19. Accelerating CNN inference on FPGAs: A Survey

    Kamel Abdelouahab, Maxime Pelcat, Jocelyn Serot +1

    cs.DCcs.ARcs.CVarXiv:1806.01683v12018
  20. What Makes Adversarial Examples Transfer Across Deepfake Detectors?

    Rafael M. Mamede, Pedro C. Neto, Ana F. Sequeira

    cs.CVcs.AIcs.CRarXiv:2609.10002v12026
  21. Albedo Estimation via Latent Bridge Matching

    Carme Corbi, David Serrano-Lozano, Javier Vazquez-Corral +1

    cs.CVcs.AIarXiv:2609.09884v12026
  22. Analysis Operator Learning and Its Application to Image Reconstruction

    Simon Hawe, Martin Kleinsteuber, Klaus Diepold

    cs.LGcs.CVarXiv:1204.5309v32012
  23. BANet: Blur-aware Attention Networks for Dynamic Scene Deblurring

    Fu-Jen Tsai, Yan-Tsung Peng, Yen-Yu Lin +2

    cs.CVarXiv:2101.07518v42021
  24. Hard negative examples are hard, but useful

    Hong Xuan, Abby Stylianou, Xiaotong Liu +1

    cs.CVcs.LGstat.MLarXiv:2007.12749v22020
  25. InverseForm: A Loss Function for Structured Boundary-Aware Segmentation

    Shubhankar Borse, Ying Wang, Yizhe Zhang +1

    cs.CVcs.LGarXiv:2104.02745v22021
  26. On the Connection between Local Attention and Dynamic Depth-wise Convolution

    Qi Han, Zejia Fan, Qi Dai +4

    cs.CVarXiv:2106.04263v52021
  27. Deep Attentive Features for Prostate Segmentation in 3D Transrectal Ultrasound

    Yi Wang, Haoran Dou, Xiaowei Hu +7

    eess.IVcs.AIcs.CVarXiv:1907.01743v22019
  28. BiHMP-GAN: Bidirectional 3D Human Motion Prediction GAN

    Jogendra Nath Kundu, Maharshi Gor, R. Venkatesh Babu

    cs.CVarXiv:1812.02591v12018
  29. FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models

    Yansen Han, Shengyi Liao, Peng Sun +4

    stat.MLcs.AIcs.CVarXiv:2609.09905v12026
  30. Reinforcement Learning with Action-Free Pre-Training from Videos

    Younggyo Seo, Kimin Lee, Stephen James +1

    cs.CVcs.AIarXiv:2203.13880v22022
  31. Strangers to Themselves: What Language Models Say About Themselves Is Generic

    Phil Blandfort, Urja Pawar

    cs.LGcs.AIcs.CLarXiv:2609.09899v12026
  32. One-shot domain adaptation in multiple sclerosis lesion segmentation using convolutional neural networks

    Sergi Valverde, Mostafa Salem, Mariano Cabezas +7

    cs.CVarXiv:1805.12415v12018
  33. Don't Forget The Past: Recurrent Depth Estimation from Monocular Video

    Vaishakh Patil, Wouter Van Gansbeke, Dengxin Dai +1

    cs.CVcs.LGcs.ROarXiv:2001.02613v22020
  34. A Generic First-Order Algorithmic Framework for Bi-Level Programming Beyond Lower-Level Singleton

    Risheng Liu, Pan Mu, Xiaoming Yuan +2

    cs.LGcs.CVmath.DSarXiv:2006.04045v22020
  35. Enhanced U-Net: A Feature Enhancement Network for Polyp Segmentation

    Krushi Patel, Andres M. Bur, Guanghui Wang

    eess.IVcs.CVarXiv:2105.00999v12021
  36. Multimodal Task-Driven Dictionary Learning for Image Classification

    Soheil Bahrampour, Nasser M. Nasrabadi, Asok Ray +1

    stat.MLcs.CVcs.LGarXiv:1502.01094v22015
  37. Generating Visual Representations for Zero-Shot Classification

    Maxime Bucher, Stéphane Herbin, Frédéric Jurie

    cs.CVcs.AIcs.LGarXiv:1708.06975v32017
  38. RangeViT: Towards Vision Transformers for 3D Semantic Segmentation in Autonomous Driving

    Angelika Ando, Spyros Gidaris, Andrei Bursuc +3

    cs.CVcs.AIcs.LGarXiv:2301.10222v22023
  39. Image Colorization with Generative Adversarial Networks

    Kamyar Nazeri, Eric Ng, Mehran Ebrahimi

    cs.CVarXiv:1803.05400v52018
  40. Block-Sparse Recovery via Convex Optimization

    Ehsan Elhamifar, Rene Vidal

    math.OCcs.CVcs.ITarXiv:1104.0654v32011
  41. Knowledge Guided Disambiguation for Large-Scale Scene Classification with Multi-Resolution CNNs

    Limin Wang, Sheng Guo, Weilin Huang +2

    cs.CVarXiv:1610.01119v22016
  42. Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational Autoencoders

    Jing Li, Di Kang, Wenjie Pei +4

    cs.CVarXiv:2108.06720v12021
  43. Event Based, Near Eye Gaze Tracking Beyond 10,000Hz

    Anastasios N. Angelopoulos, Julien N. P. Martel, Amit P. S. Kohli +2

    cs.CVcs.HCarXiv:2004.03577v32020
  44. Generation of 3D Brain MRI Using Auto-Encoding Generative Adversarial Networks

    Gihyun Kwon, Chihye Han, Dae-shik Kim

    eess.IVcs.CVarXiv:1908.02498v12019
  45. LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios

    Hanjing Zhou, Mingze Yin, Ying Lian +3

    cs.CVcs.AIcs.CLarXiv:2609.09790v12026
  46. Advancing Multimodal Medical Capabilities of Gemini

    Lin Yang, Shawn Xu, Andrew Sellergren +44

    cs.CVcs.AIcs.CLarXiv:2405.03162v12024
  47. Long-Term Human Motion Prediction by Modeling Motion Context and Enhancing Motion Dynamic

    Yongyi Tang, Lin Ma, Wei Liu +1

    cs.CVarXiv:1805.02513v12018
  48. Self-supervised Domain Adaptation for Computer Vision Tasks

    Jiaolong Xu, Liang Xiao, Antonio M. Lopez

    cs.CVcs.LGarXiv:1907.10915v32019
  49. Self-Supervised Monocular Depth Estimation with Internal Feature Fusion

    Hang Zhou, David Greenwood, Sarah Taylor

    cs.CVarXiv:2110.09482v32021
  50. A Bayesian Perspective on the Deep Image Prior

    Zezhou Cheng, Matheus Gadelha, Subhransu Maji +1

    cs.CVcs.LGstat.MLarXiv:1904.07457v12019
  51. Unsupervised Monocular Depth Learning in Dynamic Scenes

    Hanhan Li, Ariel Gordon, Hang Zhao +2

    cs.CVcs.GRcs.LGarXiv:2010.16404v22020
  52. VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

    Hu Xu, Gargi Ghosh, Po-Yao Huang +5

    cs.CVcs.CLarXiv:2105.09996v32021
  53. Distilling Image Prototypes for Guided Test-Time Adaptation

    Liwen Wang, Xingbo Dong, Iman Yi Liao +4

    cs.CVcs.AIarXiv:2609.09737v12026
  54. Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose

    Ran Yi, Zipeng Ye, Juyong Zhang +2

    cs.CVcs.GRarXiv:2002.10137v22020
  55. Error-Bounded Correction of Noisy Labels

    Songzhu Zheng, Pengxiang Wu, Aman Goswami +3

    cs.CVcs.LGarXiv:2011.10077v12020
  56. Approximate Convex Decomposition for 3D Meshes with Collision-Aware Concavity and Tree Search

    Xinyue Wei, Minghua Liu, Zhan Ling +1

    cs.GRcs.CGcs.CVarXiv:2205.02961v12022
  57. ReCo: Retrieve and Co-segment for Zero-shot Transfer

    Gyungin Shin, Weidi Xie, Samuel Albanie

    cs.CVcs.AIcs.LGarXiv:2206.07045v12022
  58. Fast, Diverse and Accurate Image Captioning Guided By Part-of-Speech

    Aditya Deshpande, Jyoti Aneja, Liwei Wang +2

    cs.CVarXiv:1805.12589v32018
  59. Ground-aware Monocular 3D Object Detection for Autonomous Driving

    Yuxuan Liu, Yuan Yixuan, Ming Liu

    cs.CVcs.ROarXiv:2102.00690v12021
  60. Object Tracking by Jointly Exploiting Frame and Event Domain

    Jiqing Zhang, Xin Yang, Yingkai Fu +3

    cs.CVarXiv:2109.09052v12021