Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,621 to 1,680 of 18,782

  1. Space-time Mixing Attention for Video Transformer

    Adrian Bulat, Juan-Manuel Perez-Rua, Swathikiran Sudhakaran +2

    cs.CVcs.AIcs.LGarXiv:2106.05968v22021
  2. TINYCD: A (Not So) Deep Learning Model For Change Detection

    Andrea Codegoni, Gabriele Lombardi, Alessandro Ferrari

    cs.CVcs.LGeess.IVarXiv:2207.13159v22022
  3. Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D Reconstruction

    Bo Yang, Sen Wang, Andrew Markham +1

    cs.CVcs.AIcs.LGarXiv:1808.00758v22018
  4. Robot Navigation in Crowds by Graph Convolutional Networks with Attention Learned from Human Gaze

    Yuying Chen, Congcong Liu, Ming Liu +1

    cs.ROcs.AIcs.CVarXiv:1909.10400v12019
  5. VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

    Jiazheng Xu, Yu Huang, Jiale Cheng +19

    cs.CVarXiv:2412.21059v42024
  6. DeepDRR -- A Catalyst for Machine Learning in Fluoroscopy-guided Procedures

    Mathias Unberath, Jan-Nico Zaech, Sing Chun Lee +4

    physics.med-phcs.CVarXiv:1803.08606v12018
  7. StopThePop: Sorted Gaussian Splatting for View-Consistent Real-time Rendering

    Lukas Radl, Michael Steiner, Mathias Parger +3

    cs.GRcs.CVarXiv:2402.00525v32024
  8. Training CNNs with Low-Rank Filters for Efficient Image Classification

    Yani Ioannou, Duncan Robertson, Jamie Shotton +2

    cs.CVcs.LGcs.NEarXiv:1511.06744v32015
  9. Face Recognition Using Deep Multi-Pose Representations

    Wael AbdAlmageed, Yue Wua, Stephen Rawlsa +9

    cs.CVarXiv:1603.07388v12016
  10. PreDiff: Precipitation Nowcasting with Latent Diffusion Models

    Zhihan Gao, Xingjian Shi, Boran Han +6

    cs.LGcs.AIcs.CVarXiv:2307.10422v22023
  11. CubiCasa5K: A Dataset and an Improved Multi-Task Model for Floorplan Image Analysis

    Ahti Kalervo, Juha Ylioinas, Markus Häikiö +2

    cs.CVarXiv:1904.01920v12019
  12. TEACHTEXT: CrossModal Generalized Distillation for Text-Video Retrieval

    Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu +4

    cs.CVarXiv:2104.08271v22021
  13. A multilevel thresholding algorithm using Electromagnetism Optimization

    Diego Oliva, Erik Cuevas, Gonzalo Pajares +2

    cs.CVarXiv:1406.6336v12014
  14. Novel Visual Category Discovery with Dual Ranking Statistics and Mutual Knowledge Distillation

    Bingchen Zhao, Kai Han

    cs.CVarXiv:2107.03358v22021
  15. MSeg3D: Multi-modal 3D Semantic Segmentation for Autonomous Driving

    Jiale Li, Hang Dai, Hao Han +1

    cs.CVarXiv:2303.08600v12023
  16. Beyond Physical Connections: Tree Models in Human Pose Estimation

    Fang Wang, Yi Li

    cs.CVarXiv:1305.2269v12013
  17. Multi-Angle Point Cloud-VAE: Unsupervised Feature Learning for 3D Point Clouds from Multiple Angles by Joint Self-Reconstruction and Half-to-Half Prediction

    Zhizhong Han, Xiyang Wang, Yu-Shen Liu +1

    cs.CVarXiv:1907.12704v12019
  18. Planar Prior Assisted PatchMatch Multi-View Stereo

    Qingshan Xu, Wenbing Tao

    cs.CVarXiv:1912.11744v12019
  19. Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

    Homanga Bharadhwaj, Roozbeh Mottaghi, Abhinav Gupta +1

    cs.ROcs.CVarXiv:2405.01527v22024
  20. A Closer Look at the Explainability of Contrastive Language-Image Pre-training

    Yi Li, Hualiang Wang, Yiqun Duan +2

    cs.CVarXiv:2304.05653v22023
  21. Total Denoising: Unsupervised Learning of 3D Point Cloud Cleaning

    Pedro Hermosilla, Tobias Ritschel, Timo Ropinski

    cs.CVcs.GRarXiv:1904.07615v22019
  22. Analyzing and Mitigating the Impact of Permanent Faults on a Systolic Array Based Neural Network Accelerator

    Jeff Zhang, Tianyu Gu, Kanad Basu +1

    cs.LGcs.ARcs.CVarXiv:1802.04657v22018
  23. VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding

    Xiang Li, Jian Ding, Mohamed Elhoseiny

    cs.CVarXiv:2406.12384v22024
  24. Out-of-Distribution Detection for Generalized Zero-Shot Action Recognition

    Devraj Mandal, Sanath Narayan, Saikumar Dwivedi +4

    cs.CVarXiv:1904.08703v22019
  25. Improved Techniques for Training Adaptive Deep Networks

    Hao Li, Hong Zhang, Xiaojuan Qi +2

    cs.CVarXiv:1908.06294v12019
  26. Avatars Grow Legs: Generating Smooth Human Motion from Sparse Tracking Inputs with Diffusion Model

    Yuming Du, Robin Kips, Albert Pumarola +3

    cs.CVarXiv:2304.08577v12023
  27. Low-Resolution Face Recognition

    Zhiyi Cheng, Xiatian Zhu, Shaogang Gong

    cs.CVarXiv:1811.08965v22018
  28. GlyphControl: Glyph Conditional Control for Visual Text Generation

    Yukang Yang, Dongnan Gui, Yuhui Yuan +4

    cs.CVarXiv:2305.18259v22023
  29. CLIP-Count: Towards Text-Guided Zero-Shot Object Counting

    Ruixiang Jiang, Lingbo Liu, Changwen Chen

    cs.CVcs.AIarXiv:2305.07304v22023
  30. Filmy Cloud Removal on Satellite Imagery with Multispectral Conditional Generative Adversarial Nets

    Kenji Enomoto, Ken Sakurada, Weimin Wang +4

    cs.CVarXiv:1710.04835v12017
  31. CovidAID: COVID-19 Detection Using Chest X-Ray

    Arpan Mangal, Surya Kalia, Harish Rajgopal +4

    eess.IVcs.CVcs.LGarXiv:2004.09803v12020
  32. Joint-task Self-supervised Learning for Temporal Correspondence

    Xueting Li, Sifei Liu, Shalini De Mello +3

    cs.CVarXiv:1909.11895v12019
  33. Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching

    Di Hu, Rui Qian, Minyue Jiang +5

    cs.CVcs.LGcs.MMarXiv:2010.05466v12020
  34. FACIAL: Synthesizing Dynamic Talking Face with Implicit Attribute Learning

    Chenxu Zhang, Yifan Zhao, Yifei Huang +4

    cs.CVarXiv:2108.07938v12021
  35. GRiT: A Generative Region-to-text Transformer for Object Understanding

    Jialian Wu, Jianfeng Wang, Zhengyuan Yang +4

    cs.CVarXiv:2212.00280v12022
  36. View-Structured Conformal Prediction for 3D Gaussian Splatting

    Junzheng Chu, Bin Pan, Zhenwei Shi

    cs.LGcs.CVarXiv:2609.10307v12026
  37. Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance

    Phuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis +4

    cs.CVarXiv:2312.10671v32023
  38. Learning A Single Network for Scale-Arbitrary Super-Resolution

    Longguang Wang, Yingqian Wang, Zaiping Lin +3

    cs.CVarXiv:2004.03791v22020
  39. GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

    Drew A. Hudson, Christopher D. Manning

    cs.CLcs.AIcs.CVarXiv:1902.09506v32019
  40. Point-Set Anchors for Object Detection, Instance Segmentation and Pose Estimation

    Fangyun Wei, Xiao Sun, Hongyang Li +2

    cs.CVarXiv:2007.02846v42020
  41. Towards Flexible Blind JPEG Artifacts Removal

    Jiaxi Jiang, Kai Zhang, Radu Timofte

    eess.IVcs.CVarXiv:2109.14573v12021
  42. SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition

    Mingxin Huang, Yuliang Liu, Zhenghao Peng +6

    cs.CVarXiv:2203.10209v12022
  43. Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code

    Xuan Ju, Ailing Zeng, Yuxuan Bian +2

    cs.CVarXiv:2310.01506v22023
  44. Text-Only Training for Image Captioning using Noise-Injected CLIP

    David Nukrai, Ron Mokady, Amir Globerson

    cs.CVcs.AIcs.LGarXiv:2211.00575v12022
  45. Nutrition5k: Towards Automatic Nutritional Understanding of Generic Food

    Quin Thames, Arjun Karpur, Wade Norris +4

    cs.CVcs.LGarXiv:2103.03375v22021
  46. Feature Learning from Incomplete EEG with Denoising Autoencoder

    Junhua Li, Zbigniew Struzik, Liqing Zhang +1

    cs.CVq-bio.NCarXiv:1410.0818v12014
  47. Uncertainty-Informed Deep Learning Models Enable High-Confidence Predictions for Digital Histopathology

    James M Dolezal, Andrew Srisuwananukorn, Dmitry Karpeyev +13

    q-bio.QMcs.CVeess.IVarXiv:2204.04516v12022
  48. Curriculum Model Adaptation with Synthetic and Real Data for Semantic Foggy Scene Understanding

    Dengxin Dai, Christos Sakaridis, Simon Hecker +1

    cs.CVarXiv:1901.01415v22019
  49. Cross-Attention of Disentangled Modalities for 3D Human Mesh Recovery with Transformers

    Junhyeong Cho, Kim Youwang, Tae-Hyun Oh

    cs.CVcs.AIcs.LGarXiv:2207.13820v12022
  50. MMM: Generative Masked Motion Model

    Ekkasit Pinyoanuntapong, Pu Wang, Minwoo Lee +1

    cs.CVcs.AIcs.LGarXiv:2312.03596v22023
  51. DigiFace-1M: 1 Million Digital Face Images for Face Recognition

    Gwangbin Bae, Martin de La Gorce, Tadas Baltrusaitis +5

    cs.CVarXiv:2210.02579v12022
  52. Hashing on Nonlinear Manifolds

    Fumin Shen, Chunhua Shen, Qinfeng Shi +3

    cs.CVarXiv:1412.0826v12014
  53. Manipulation by Feel: Touch-Based Control with Deep Predictive Models

    Stephen Tian, Frederik Ebert, Dinesh Jayaraman +4

    cs.ROcs.AIcs.CVarXiv:1903.04128v12019
  54. Dark, Beyond Deep: A Paradigm Shift to Cognitive AI with Humanlike Common Sense

    Yixin Zhu, Tao Gao, Lifeng Fan +9

    cs.AIcs.CVcs.LGarXiv:2004.09044v12020
  55. Photorealistic Image Synthesis for Object Instance Detection

    Tomas Hodan, Vibhav Vineet, Ran Gal +6

    cs.CVcs.AIcs.ROarXiv:1902.03334v12019
  56. Temporal Pyramid Pooling Based Convolutional Neural Networks for Action Recognition

    Peng Wang, Yuanzhouhan Cao, Chunhua Shen +2

    cs.CVarXiv:1503.01224v22015
  57. Controlling Vision-Language Models for Multi-Task Image Restoration

    Ziwei Luo, Fredrik K. Gustafsson, Zheng Zhao +2

    cs.CVarXiv:2310.01018v22023
  58. Active Learning for Deep Detection Neural Networks

    Hamed H. Aghdam, Abel Gonzalez-Garcia, Joost van de Weijer +1

    cs.CVcs.LGarXiv:1911.09168v12019
  59. Unsupervised Domain Adaptation via Disentangled Representations: Application to Cross-Modality Liver Segmentation

    Junlin Yang, Nicha C. Dvornek, Fan Zhang +3

    eess.IVcs.CVarXiv:1907.13590v22019
  60. Newtonian Image Understanding: Unfolding the Dynamics of Objects in Static Images

    Roozbeh Mottaghi, Hessam Bagherinezhad, Mohammad Rastegari +1

    cs.CVarXiv:1511.04048v12015