Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,441 to 7,500 of 18,957

  1. HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

    Mude Hui, Siwei Yang, Bingchen Zhao +5

    cs.CVcs.AIarXiv:2404.09990v12024
  2. Sample4Geo: Hard Negative Sampling For Cross-View Geo-Localisation

    Fabian Deuser, Konrad Habel, Norbert Oswald

    cs.CVarXiv:2303.11851v22023
  3. Flow-edge Guided Video Completion

    Chen Gao, Ayush Saraf, Jia-Bin Huang +1

    cs.CVarXiv:2009.01835v12020
  4. VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

    Xinhao Li, Yi Wang, Jiashuo Yu +10

    cs.CVcs.LGarXiv:2501.00574v42024
  5. Autoregressive Queries for Adaptive Tracking with Spatio-TemporalTransformers

    Jinxia Xie, Bineng Zhong, Zhiyi Mo +4

    cs.CVarXiv:2403.10574v12024
  6. Cross-Domain Gradient Discrepancy Minimization for Unsupervised Domain Adaptation

    Zhekai Du, Jingjing Li, Hongzu Su +2

    cs.CVcs.AIcs.LGarXiv:2106.04151v12021
  7. GBDT-MO: Gradient Boosted Decision Trees for Multiple Outputs

    Zhendong Zhang, Cheolkon Jung

    cs.CVcs.LGarXiv:1909.04373v22019
  8. DMT: Dynamic Mutual Training for Semi-Supervised Learning

    Zhengyang Feng, Qianyu Zhou, Qiqi Gu +5

    cs.CVarXiv:2004.08514v42020
  9. Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

    Yuexiang Zhai, Hao Bai, Zipeng Lin +8

    cs.AIcs.CLcs.CVarXiv:2405.10292v32024
  10. ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts

    Mu Cai, Haotian Liu, Dennis Park +4

    cs.CVcs.AIcs.CLarXiv:2312.00784v22023
  11. A Lightweight Optical Flow CNN - Revisiting Data Fidelity and Regularization

    Tak-Wai Hui, Xiaoou Tang, Chen Change Loy

    cs.CVarXiv:1903.07414v32019
  12. Generative models improve fairness of medical classifiers under distribution shifts

    Ira Ktena, Olivia Wiles, Isabela Albuquerque +9

    cs.CVarXiv:2304.09218v12023
  13. Multi-Scale 3D Gaussian Splatting for Anti-Aliased Rendering

    Zhiwen Yan, Weng Fei Low, Yu Chen +1

    cs.CVarXiv:2311.17089v22023
  14. Waste detection in Pomerania: non-profit project for detecting waste in environment

    Sylwia Majchrowska, Agnieszka Mikołajczyk, Maria Ferlin +4

    cs.CVeess.IVarXiv:2105.06808v12021
  15. RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation

    Xiaolei Lang, Ze Kang, Zehao Huang +1

    cs.CVarXiv:2609.02847v12026
  16. A 3D Probabilistic Deep Learning System for Detection and Diagnosis of Lung Cancer Using Low-Dose CT Scans

    Onur Ozdemir, Rebecca L. Russell, Andrew A. Berlin

    cs.CVcs.LGarXiv:1902.03233v32019
  17. VIPS: Vehicle-Infrastructure Cooperative Planning Benchmark via Pseudo-Simulation

    Hoonhee Cho, Jae-Young Kang, Giwon Lee +3

    cs.CVarXiv:2609.02462v12026
  18. Vision-Aided 6G Wireless Communications: Blockage Prediction and Proactive Handoff

    Gouranga Charan, Muhammad Alrabeiah, Ahmed Alkhateeb

    eess.SPcs.CVarXiv:2102.09527v22021
  19. Talking-head Generation with Rhythmic Head Motion

    Lele Chen, Guofeng Cui, Celong Liu +4

    cs.CVcs.GRarXiv:2007.08547v12020
  20. Low-resolution Face Recognition in the Wild via Selective Knowledge Distillation

    Shiming Ge, Shengwei Zhao, Chenyu Li +1

    cs.CVarXiv:1811.09998v22018
  21. InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models

    Chengyin Hu, Dingyi Lu, Jiaju Han +5

    cs.CVcs.AIarXiv:2609.02233v12026
  22. Automatic Image Segmentation by Dynamic Region Merging

    Bo Peng, Lei Zhang, David Zhang

    cs.CVcs.ROarXiv:1012.1193v12010
  23. ZipIt! Merging Models from Different Tasks without Training

    George Stoica, Daniel Bolya, Jakob Bjorner +3

    cs.CVcs.LGarXiv:2305.03053v32023
  24. Learning Analysis-by-Synthesis for 6D Pose Estimation in RGB-D Images

    Alexander Krull, Eric Brachmann, Frank Michel +3

    cs.CVarXiv:1508.04546v12015
  25. Dense Human Body Correspondences Using Convolutional Networks

    Lingyu Wei, Qixing Huang, Duygu Ceylan +2

    cs.CVcs.GRarXiv:1511.05904v22015
  26. PolyLoss: A Polynomial Expansion Perspective of Classification Loss Functions

    Zhaoqi Leng, Mingxing Tan, Chenxi Liu +4

    cs.CVarXiv:2204.12511v22022
  27. Differentiable Prompt Makes Pre-trained Language Models Better Few-shot Learners

    Ningyu Zhang, Luoqiu Li, Xiang Chen +5

    cs.CLcs.AIcs.CVarXiv:2108.13161v72021
  28. VisualMRC: Machine Reading Comprehension on Document Images

    Ryota Tanaka, Kyosuke Nishida, Sen Yoshida

    cs.CLcs.CVarXiv:2101.11272v22021
  29. Exploiting saliency for object segmentation from image level labels

    Seong Joon Oh, Rodrigo Benenson, Anna Khoreva +3

    cs.CVarXiv:1701.08261v22017
  30. CNN-based Lidar Point Cloud De-Noising in Adverse Weather

    Robin Heinzler, Florian Piewak, Philipp Schindler +1

    cs.CVcs.ROarXiv:1912.03874v22019
  31. The impact of patient clinical information on automated skin cancer detection

    Andre G. C. Pacheco, Renato A. Krohling

    eess.IVcs.CVcs.LGarXiv:1909.12912v12019
  32. Remember Intentions: Retrospective-Memory-based Trajectory Prediction

    Chenxin Xu, Weibo Mao, Wenjun Zhang +1

    cs.CVarXiv:2203.11474v12022
  33. An interpretable classifier for high-resolution breast cancer screening images utilizing weakly supervised localization

    Yiqiu Shen, Nan Wu, Jason Phang +8

    cs.CVcs.LGeess.IVarXiv:2002.07613v12020
  34. Channel Gating Neural Networks

    Weizhe Hua, Yuan Zhou, Christopher De Sa +2

    cs.LGcs.CVstat.MLarXiv:1805.12549v22018
  35. EmergencyNet: Efficient Aerial Image Classification for Drone-Based Emergency Monitoring Using Atrous Convolutional Feature Fusion

    Christos Kyrkou, Theocharis Theocharides

    cs.CVcs.LGcs.ROarXiv:2104.14006v12021
  36. Differentially Private Paired Table-Image Multimodal Synthesis

    Kai Chen, Josephine Lamp, Somesh Jha +1

    cs.CRcs.AIcs.CVarXiv:2609.00708v12026
  37. Diversity with Cooperation: Ensemble Methods for Few-Shot Classification

    Nikita Dvornik, Cordelia Schmid, Julien Mairal

    cs.CVcs.AIarXiv:1903.11341v22019
  38. ALIKED: A Lighter Keypoint and Descriptor Extraction Network via Deformable Transformation

    Xiaoming Zhao, Xingming Wu, Weihai Chen +3

    cs.CVarXiv:2304.03608v22023
  39. Frame attention networks for facial expression recognition in videos

    Debin Meng, Xiaojiang Peng, Kai Wang +1

    cs.CVcs.HCcs.MMarXiv:1907.00193v22019
  40. Image-based Localization using Hourglass Networks

    Iaroslav Melekhov, Juha Ylioinas, Juho Kannala +1

    cs.CVarXiv:1703.07971v32017
  41. MineGAN: effective knowledge transfer from GANs to target domains with few images

    Yaxing Wang, Abel Gonzalez-Garcia, David Berga +3

    cs.CVarXiv:1912.05270v32019
  42. Multi-scale recognition with DAG-CNNs

    Songfan Yang, Deva Ramanan

    cs.CVarXiv:1505.05232v12015
  43. Unpaired Motion Style Transfer from Video to Animation

    Kfir Aberman, Yijia Weng, Dani Lischinski +2

    cs.GRcs.CVcs.LGarXiv:2005.05751v12020
  44. CNN in MRF: Video Object Segmentation via Inference in A CNN-Based Higher-Order Spatio-Temporal MRF

    Linchao Bao, Baoyuan Wu, Wei Liu

    cs.CVarXiv:1803.09453v12018
  45. DAiSEE: Towards User Engagement Recognition in the Wild

    Abhay Gupta, Arjun D'Cunha, Kamal Awasthi +1

    cs.CVcs.LGarXiv:1609.01885v72016
  46. MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph Captioning

    Jie Lei, Liwei Wang, Yelong Shen +3

    cs.CLcs.CVcs.LGarXiv:2005.05402v12020
  47. One-Shot Neural Architecture Search via Self-Evaluated Template Network

    Xuanyi Dong, Yi Yang

    cs.CVarXiv:1910.05733v42019
  48. Exploiting BERT For Multimodal Target Sentiment Classification Through Input Space Translation

    Zaid Khan, Yun Fu

    cs.CLcs.CVarXiv:2108.01682v22021
  49. Graph-Based Object Classification for Neuromorphic Vision Sensing

    Yin Bi, Aaron Chadha, Alhabib Abbas +2

    cs.CVarXiv:1908.06648v12019
  50. Less Is More: Picking Informative Frames for Video Captioning

    Yangyu Chen, Shuhui Wang, Weigang Zhang +1

    cs.CVarXiv:1803.01457v12018
  51. Explainable Neural Computation via Stack Neural Module Networks

    Ronghang Hu, Jacob Andreas, Trevor Darrell +1

    cs.CVarXiv:1807.08556v32018
  52. Doppio: A Dataset for Contactless Weight Estimation of Falling Particles

    Simon Kiefhaber, Jan-Martin O. Steitz, Julia Grabinski +5

    cs.CVarXiv:2609.02528v12026
  53. Applications of Artificial Neural Networks in Microorganism Image Analysis: A Comprehensive Review from Conventional Multilayer Perceptron to Popular Convolutional Neural Network and Potential Visual Transformer

    Jinghua Zhang, Chen Li, Yimin Yin +2

    cs.CVcs.AIarXiv:2108.00358v32021
  54. Artificial Intelligence-Based Methods for Fusion of Electronic Health Records and Imaging Data

    Farida Mohsen, Hazrat Ali, Nady El Hajj +1

    cs.LGcs.AIcs.CVarXiv:2210.13462v12022
  55. Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap

    Nirajan Kunwor, Sanjaya Poudel, Quoc-Huy Trinh +2

    cs.CVcs.AIcs.LGarXiv:2609.02111v12026
  56. Mamba YOLO: A Simple Baseline for Object Detection with State Space Model

    Zeyu Wang, Chen Li, Huiying Xu +2

    cs.CVarXiv:2406.05835v22024
  57. SPA-GAN: Spatial Attention GAN for Image-to-Image Translation

    Hajar Emami, Majid Moradi Aliabadi, Ming Dong +1

    cs.CVarXiv:1908.06616v32019
  58. Deep Learning-based 3D Point Cloud Classification: A Systematic Survey and Outlook

    Huang Zhang, Changshuo Wang, Shengwei Tian +4

    cs.CVarXiv:2311.02608v12023
  59. Local contrastive loss with pseudo-label based self-training for semi-supervised medical image segmentation

    Krishna Chaitanya, Ertunc Erdil, Neerav Karani +1

    cs.CVcs.AIcs.LGarXiv:2112.09645v12021
  60. Attention-Aware Face Hallucination via Deep Reinforcement Learning

    Qingxing Cao, Liang Lin, Yukai Shi +2

    cs.CVarXiv:1708.03132v12017