Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,561 to 1,620 of 18,779

  1. MoVQ: Modulating Quantized Vectors for High-Fidelity Image Generation

    Chuanxia Zheng, Long Tung Vuong, Jianfei Cai +1

    cs.CVarXiv:2209.09002v12022
  2. Active Transfer Learning Network: A Unified Deep Joint Spectral-Spatial Feature Learning Model For Hyperspectral Image Classification

    Cheng Deng, Yumeng Xue, Xianglong Liu +2

    cs.CVarXiv:1904.02454v12019
  3. Learning Blind Motion Deblurring

    Patrick Wieschollek, Michael Hirsch, Bernhard Schölkopf +1

    cs.CVarXiv:1708.04208v12017
  4. Pedestrian Path, Pose and Intention Prediction through Gaussian Process Dynamical Models and Pedestrian Activity Recognition

    Raul Quintero, Ignacio Parra, David Fernandez Llorca +1

    cs.CVarXiv:2004.14747v12020
  5. A deep learning approach to detecting volcano deformation from satellite imagery using synthetic datasets

    Nantheera Anantrasirichai, Juliet Biggs, Fabien Albino +1

    cs.CVeess.IVarXiv:1905.07286v12019
  6. COUCH: Towards Controllable Human-Chair Interactions

    Xiaohan Zhang, Bharat Lal Bhatnagar, Vladimir Guzov +2

    cs.CVarXiv:2205.00541v12022
  7. Visual Search at Pinterest

    Yushi Jing, David Liu, Dmitry Kislyuk +4

    cs.CVarXiv:1505.07647v32015
  8. Deep Learning in Diabetic Foot Ulcers Detection: A Comprehensive Evaluation

    Moi Hoon Yap, Ryo Hachiuma, Azadeh Alavi +18

    cs.CVarXiv:2010.03341v32020
  9. LightMedSeg-ISLES: Stroke Lesion Segmentation with 81x Fewer Parameters than nnU-Net

    Giorgi Nikvashvili, Hanxue Gu, Jie Bao +2

    cs.CVcs.LGarXiv:2609.09634v12026
  10. Discuss Before Moving: Visual Language Navigation via Multi-expert Discussions

    Yuxing Long, Xiaoqi Li, Wenzhe Cai +1

    cs.ROcs.AIcs.CLarXiv:2309.11382v12023
  11. Probabilistic Monocular 3D Human Pose Estimation with Normalizing Flows

    Tom Wehrbein, Marco Rudolph, Bodo Rosenhahn +1

    cs.CVarXiv:2107.13788v22021
  12. ZigMa: A DiT-style Zigzag Mamba Diffusion Model

    Vincent Tao Hu, Stefan Andreas Baumann, Ming Gui +4

    cs.CVcs.AIcs.CLarXiv:2403.13802v32024
  13. MethaneFuse: Learning from Multi-Sensor Satellite Observations for Methane Plume Detection

    Yuyao Wang, Juliana Y. Leung, Di Niu

    cs.CVcs.LGarXiv:2609.09762v12026
  14. LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis

    Zejiang Shen, Ruochen Zhang, Melissa Dell +3

    cs.CVcs.AIarXiv:2103.15348v22021
  15. Meta-Learning with Task-Adaptive Loss Function for Few-Shot Learning

    Sungyong Baik, Janghoon Choi, Heewon Kim +3

    cs.LGcs.CVarXiv:2110.03909v22021
  16. Domain Generalization with Domain-Specific Aggregation Modules

    Antonio D'Innocente, Barbara Caputo

    cs.CVarXiv:1809.10966v12018
  17. Deep Hyperspherical Learning

    Weiyang Liu, Yan-Ming Zhang, Xingguo Li +4

    cs.LGcs.CVstat.MLarXiv:1711.03189v52017
  18. Fine-Grained Car Detection for Visual Census Estimation

    Timnit Gebru, Jonathan Krause, Yilun Wang +3

    cs.CVarXiv:1709.02480v12017
  19. TAC-GAN - Text Conditioned Auxiliary Classifier Generative Adversarial Network

    Ayushman Dash, John Cristian Borges Gamboa, Sheraz Ahmed +2

    cs.CVarXiv:1703.06412v22017
  20. A General Pipeline for 3D Detection of Vehicles

    Xinxin Du, Marcelo H. Ang, Sertac Karaman +1

    cs.CVeess.IVstat.MLarXiv:1803.00387v12018
  21. 2D Car Detection in Radar Data with PointNets

    Andreas Danzer, Thomas Griebel, Martin Bach +1

    cs.CVcs.LGstat.MLarXiv:1904.08414v32019
  22. D-Grasp: Physically Plausible Dynamic Grasp Synthesis for Hand-Object Interactions

    Sammy Christen, Muhammed Kocabas, Emre Aksan +3

    cs.CVcs.LGcs.ROarXiv:2112.03028v22021
  23. Top-Down Feedback for Crowd Counting Convolutional Neural Network

    Deepak Babu Sam, R. Venkatesh Babu

    cs.CVarXiv:1807.08881v22018
  24. Learning Disentangled Semantic Representation for Domain Adaptation

    Ruichu Cai, Zijian Li, Pengfei Wei +3

    cs.CVcs.LGarXiv:2012.11807v12020
  25. Retrieval Augmented Classification for Long-Tail Visual Recognition

    Alexander Long, Wei Yin, Thalaiyasingam Ajanthan +6

    cs.CVarXiv:2202.11233v12022
  26. Face Anti-Spoofing with Human Material Perception

    Zitong Yu, Xiaobai Li, Xuesong Niu +2

    cs.CVarXiv:2007.02157v12020
  27. DeepTravel: a Neural Network Based Travel Time Estimation Model with Auxiliary Supervision

    Hanyuan Zhang, Hao Wu, Weiwei Sun +1

    cs.LGcs.CVstat.MLarXiv:1802.02147v12018
  28. End-to-end Trained CNN Encode-Decoder Networks for Image Steganography

    Atique ur Rehman, Rafia Rahim, M Shahroz Nadeem +1

    cs.MMcs.CVarXiv:1711.07201v12017
  29. ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation

    Guanxing Lu, Shiyi Zhang, Ziwei Wang +3

    cs.ROcs.CVarXiv:2403.08321v22024
  30. From ImageNet to Image Classification: Contextualizing Progress on Benchmarks

    Dimitris Tsipras, Shibani Santurkar, Logan Engstrom +2

    cs.CVcs.LGstat.MLarXiv:2005.11295v12020
  31. LeCor: Learning to Be Corrected by Meta-Learned Test-Time Training for Interactive 3D Lung-Tumour Segmentation

    Yi Luo, Yike Guo, Wenxuan Li +3

    cs.CVcs.LGarXiv:2609.09477v12026
  32. All-In-One Underwater Image Enhancement using Domain-Adversarial Learning

    Pritish Uplavikar, Zhenyu Wu, Zhangyang Wang

    cs.CVarXiv:1905.13342v12019
  33. We are More than Our Joints: Predicting how 3D Bodies Move

    Yan Zhang, Michael J. Black, Siyu Tang

    cs.CVarXiv:2012.00619v22020
  34. Joint super-resolution and synthesis of 1 mm isotropic MP-RAGE volumes from clinical MRI exams with scans of different orientation, resolution and contrast

    Juan Eugenio Iglesias, Benjamin Billot, Yael Balbastre +6

    eess.IVcs.CVarXiv:2012.13340v12020
  35. Auxiliary Signal-Guided Knowledge Encoder-Decoder for Medical Report Generation

    Mingjie Li, Fuyu Wang, Xiaojun Chang +1

    cs.CVcs.CLeess.IVarXiv:2006.03744v12020
  36. 3D Human Pose Estimation via Intuitive Physics

    Shashank Tripathi, Lea Müller, Chun-Hao P. Huang +3

    cs.CVcs.AIcs.GRarXiv:2303.18246v32023
  37. KING: Generating Safety-Critical Driving Scenarios for Robust Imitation via Kinematics Gradients

    Niklas Hanselmann, Katrin Renz, Kashyap Chitta +2

    cs.ROcs.CVcs.LGarXiv:2204.13683v12022
  38. Space-time Mixing Attention for Video Transformer

    Adrian Bulat, Juan-Manuel Perez-Rua, Swathikiran Sudhakaran +2

    cs.CVcs.AIcs.LGarXiv:2106.05968v22021
  39. TINYCD: A (Not So) Deep Learning Model For Change Detection

    Andrea Codegoni, Gabriele Lombardi, Alessandro Ferrari

    cs.CVcs.LGeess.IVarXiv:2207.13159v22022
  40. Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D Reconstruction

    Bo Yang, Sen Wang, Andrew Markham +1

    cs.CVcs.AIcs.LGarXiv:1808.00758v22018
  41. Robot Navigation in Crowds by Graph Convolutional Networks with Attention Learned from Human Gaze

    Yuying Chen, Congcong Liu, Ming Liu +1

    cs.ROcs.AIcs.CVarXiv:1909.10400v12019
  42. VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

    Jiazheng Xu, Yu Huang, Jiale Cheng +19

    cs.CVarXiv:2412.21059v42024
  43. DeepDRR -- A Catalyst for Machine Learning in Fluoroscopy-guided Procedures

    Mathias Unberath, Jan-Nico Zaech, Sing Chun Lee +4

    physics.med-phcs.CVarXiv:1803.08606v12018
  44. StopThePop: Sorted Gaussian Splatting for View-Consistent Real-time Rendering

    Lukas Radl, Michael Steiner, Mathias Parger +3

    cs.GRcs.CVarXiv:2402.00525v32024
  45. Training CNNs with Low-Rank Filters for Efficient Image Classification

    Yani Ioannou, Duncan Robertson, Jamie Shotton +2

    cs.CVcs.LGcs.NEarXiv:1511.06744v32015
  46. Face Recognition Using Deep Multi-Pose Representations

    Wael AbdAlmageed, Yue Wua, Stephen Rawlsa +9

    cs.CVarXiv:1603.07388v12016
  47. PreDiff: Precipitation Nowcasting with Latent Diffusion Models

    Zhihan Gao, Xingjian Shi, Boran Han +6

    cs.LGcs.AIcs.CVarXiv:2307.10422v22023
  48. CubiCasa5K: A Dataset and an Improved Multi-Task Model for Floorplan Image Analysis

    Ahti Kalervo, Juha Ylioinas, Markus Häikiö +2

    cs.CVarXiv:1904.01920v12019
  49. TEACHTEXT: CrossModal Generalized Distillation for Text-Video Retrieval

    Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu +4

    cs.CVarXiv:2104.08271v22021
  50. A multilevel thresholding algorithm using Electromagnetism Optimization

    Diego Oliva, Erik Cuevas, Gonzalo Pajares +2

    cs.CVarXiv:1406.6336v12014
  51. Novel Visual Category Discovery with Dual Ranking Statistics and Mutual Knowledge Distillation

    Bingchen Zhao, Kai Han

    cs.CVarXiv:2107.03358v22021
  52. MSeg3D: Multi-modal 3D Semantic Segmentation for Autonomous Driving

    Jiale Li, Hang Dai, Hao Han +1

    cs.CVarXiv:2303.08600v12023
  53. Beyond Physical Connections: Tree Models in Human Pose Estimation

    Fang Wang, Yi Li

    cs.CVarXiv:1305.2269v12013
  54. Multi-Angle Point Cloud-VAE: Unsupervised Feature Learning for 3D Point Clouds from Multiple Angles by Joint Self-Reconstruction and Half-to-Half Prediction

    Zhizhong Han, Xiyang Wang, Yu-Shen Liu +1

    cs.CVarXiv:1907.12704v12019
  55. Planar Prior Assisted PatchMatch Multi-View Stereo

    Qingshan Xu, Wenbing Tao

    cs.CVarXiv:1912.11744v12019
  56. Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

    Homanga Bharadhwaj, Roozbeh Mottaghi, Abhinav Gupta +1

    cs.ROcs.CVarXiv:2405.01527v22024
  57. A Closer Look at the Explainability of Contrastive Language-Image Pre-training

    Yi Li, Hualiang Wang, Yiqun Duan +2

    cs.CVarXiv:2304.05653v22023
  58. Total Denoising: Unsupervised Learning of 3D Point Cloud Cleaning

    Pedro Hermosilla, Tobias Ritschel, Timo Ropinski

    cs.CVcs.GRarXiv:1904.07615v22019
  59. Analyzing and Mitigating the Impact of Permanent Faults on a Systolic Array Based Neural Network Accelerator

    Jeff Zhang, Tianyu Gu, Kanad Basu +1

    cs.LGcs.ARcs.CVarXiv:1802.04657v22018
  60. VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding

    Xiang Li, Jian Ding, Mohamed Elhoseiny

    cs.CVarXiv:2406.12384v22024