Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,921 to 13,980 of 18,830

  1. SalsaNext: Fast, Uncertainty-aware Semantic Segmentation of LiDAR Point Clouds for Autonomous Driving

    Tiago Cortinhal, George Tzelepis, Eren Erdal Aksoy

    cs.CVcs.LGarXiv:2003.03653v42020
  2. Semi-Supervised Adaptation of Vision-Language Models for Image Classification

    Mohamed L. Mekhalfi, Mohamad M. Al Rahhal, Yakoub Bazi +4

    cs.CVarXiv:2608.25485v12026
  3. MotionCLIP: Exposing Human Motion Generation to CLIP Space

    Guy Tevet, Brian Gordon, Amir Hertz +2

    cs.CVcs.GRarXiv:2203.08063v12022
  4. TD-MPC2: Scalable, Robust World Models for Continuous Control

    Nicklas Hansen, Hao Su, Xiaolong Wang

    cs.LGcs.AIcs.CVarXiv:2310.16828v22023
  5. HATS: Histograms of Averaged Time Surfaces for Robust Event-based Object Classification

    Amos Sironi, Manuele Brambilla, Nicolas Bourdis +2

    cs.CVarXiv:1803.07913v12018
  6. Convolutional Neural Networks Applied to House Numbers Digit Classification

    Pierre Sermanet, Soumith Chintala, Yann LeCun

    cs.CVcs.LGcs.NEarXiv:1204.3968v12012
  7. Large Separable Kernel Attention: Rethinking the Large Kernel Attention Design in CNN

    Kin Wai Lau, Lai-Man Po, Yasar Abbas Ur Rehman

    cs.CVarXiv:2309.01439v32023
  8. Normalized Loss Functions for Deep Learning with Noisy Labels

    Xingjun Ma, Hanxun Huang, Yisen Wang +3

    cs.LGcs.CVstat.MLarXiv:2006.13554v12020
  9. Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude Learning

    Yu Tian, Guansong Pang, Yuanhong Chen +3

    cs.CVarXiv:2101.10030v32021
  10. Privacy-preserving Federated Brain Tumour Segmentation

    Wenqi Li, Fausto Milletarì, Daguang Xu +8

    cs.CVarXiv:1910.00962v12019
  11. Suggestive Annotation: A Deep Active Learning Framework for Biomedical Image Segmentation

    Lin Yang, Yizhe Zhang, Jianxu Chen +2

    cs.CVarXiv:1706.04737v12017
  12. Few-Shot Class-Incremental Learning

    Xiaoyu Tao, Xiaopeng Hong, Xinyuan Chang +3

    cs.CVcs.LGstat.MLarXiv:2004.10956v22020
  13. Maintaining Discrimination and Fairness in Class Incremental Learning

    Bowen Zhao, Xi Xiao, Guojun Gan +2

    cs.CVarXiv:1911.07053v12019
  14. GANerated Hands for Real-time 3D Hand Tracking from Monocular RGB

    Franziska Mueller, Florian Bernard, Oleksandr Sotnychenko +4

    cs.CVarXiv:1712.01057v12017
  15. HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

    Linjie Li, Yen-Chun Chen, Yu Cheng +3

    cs.CVcs.CLcs.LGarXiv:2005.00200v22020
  16. TEA: Temporal Excitation and Aggregation for Action Recognition

    Yan Li, Bin Ji, Xintian Shi +3

    cs.CVarXiv:2004.01398v12020
  17. Fast Segment Anything

    Xu Zhao, Wenchao Ding, Yongqi An +5

    cs.CVcs.AIarXiv:2306.12156v12023
  18. Unsupervised Post-Training of Foundation Models: A Survey

    Yijie Xu, Qianyi Cai, Huizai Yao +9

    cs.CLcs.AIcs.CVarXiv:2608.24982v12026
  19. A Patient-Centric Dataset of Images and Metadata for Identifying Melanomas Using Clinical Context

    Veronica Rotemberg, Nicholas Kurtansky, Brigid Betz-Stablein +21

    eess.IVcs.CVcs.CYarXiv:2008.07360v12020
  20. Total Capture: A 3D Deformation Model for Tracking Faces, Hands, and Bodies

    Hanbyul Joo, Tomas Simon, Yaser Sheikh

    cs.CVarXiv:1801.01615v12018
  21. DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

    Xiaoyu Tian, Junru Gu, Bailin Li +7

    cs.CVarXiv:2402.12289v52024
  22. Demystifying Neural Style Transfer

    Yanghao Li, Naiyan Wang, Jiaying Liu +1

    cs.CVcs.LGcs.NEarXiv:1701.01036v22017
  23. Beyond Finite Layer Neural Networks: Bridging Deep Architectures and Numerical Differential Equations

    Yiping Lu, Aoxiao Zhong, Quanzheng Li +1

    cs.CVcs.LGstat.MLarXiv:1710.10121v32017
  24. Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language Models

    Manli Shu, Weili Nie, De-An Huang +4

    cs.CVarXiv:2209.07511v12022
  25. Repulsion Loss: Detecting Pedestrians in a Crowd

    Xinlong Wang, Tete Xiao, Yuning Jiang +3

    cs.CVarXiv:1711.07752v22017
  26. FlowNet3D: Learning Scene Flow in 3D Point Clouds

    Xingyu Liu, Charles R. Qi, Leonidas J. Guibas

    cs.CVcs.LGarXiv:1806.01411v32018
  27. Visualizing Deep Convolutional Neural Networks Using Natural Pre-Images

    Aravindh Mahendran, Andrea Vedaldi

    cs.CVarXiv:1512.02017v32015
  28. YOLO-LITE: A Real-Time Object Detection Algorithm Optimized for Non-GPU Computers

    Jonathan Pedoeem, Rachel Huang

    cs.CVarXiv:1811.05588v12018
  29. Tell Me Where to Look: Guided Attention Inference Network

    Kunpeng Li, Ziyan Wu, Kuan-Chuan Peng +2

    cs.CVcs.LGarXiv:1802.10171v12018
  30. DSEC: A Stereo Event Camera Dataset for Driving Scenarios

    Mathias Gehrig, Willem Aarents, Daniel Gehrig +1

    cs.CVcs.ROarXiv:2103.06011v12021
  31. Found in Translation: Learning Robust Joint Representations by Cyclic Translations Between Modalities

    Hai Pham, Paul Pu Liang, Thomas Manzini +2

    cs.LGcs.CLcs.CVarXiv:1812.07809v22018
  32. Detail-revealing Deep Video Super-resolution

    Xin Tao, Hongyun Gao, Renjie Liao +2

    cs.CVarXiv:1704.02738v12017
  33. A Short Note on the Kinetics-700 Human Action Dataset

    Joao Carreira, Eric Noland, Chloe Hillier +1

    cs.CVarXiv:1907.06987v22019
  34. Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

    Alexander Ku, Peter Anderson, Roma Patel +2

    cs.CVcs.AIcs.CLarXiv:2010.07954v12020
  35. NiftyNet: a deep-learning platform for medical imaging

    Eli Gibson, Wenqi Li, Carole Sudre +14

    cs.CVcs.LGcs.NEarXiv:1709.03485v22017
  36. ResViT: Residual vision transformers for multi-modal medical image synthesis

    Onat Dalmaz, Mahmut Yurt, Tolga Çukur

    eess.IVcs.CVarXiv:2106.16031v32021
  37. In-Place Scene Labelling and Understanding with Implicit Scene Representation

    Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger +1

    cs.CVarXiv:2103.15875v22021
  38. Contextrast++: Robust Multi-Scale Contextual Contrastive Learning for Semantic Segmentation

    Changki Sung, Hyungtae Lim, Wanhee Kim +2

    cs.CVcs.ROarXiv:2608.22679v12026
  39. First-Person Hand Action Benchmark with RGB-D Videos and 3D Hand Pose Annotations

    Guillermo Garcia-Hernando, Shanxin Yuan, Seungryul Baek +1

    cs.CVarXiv:1704.02463v22017
  40. Learning Trajectory Dependencies for Human Motion Prediction

    Wei Mao, Miaomiao Liu, Mathieu Salzmann +1

    cs.CVarXiv:1908.05436v32019
  41. Reliability- and Anatomy-Consistency-Aware Multimodal Learning for Robust Fracture Classification from Bangladeshi Radiographs

    Musa Tur Farazi, K G Subarno Bithi

    eess.IVcs.AIcs.CVarXiv:2608.21482v12026
  42. Convolutional Recurrent Neural Networks for Dynamic MR Image Reconstruction

    Chen Qin, Jo Schlemper, Jose Caballero +3

    cs.CVarXiv:1712.01751v32017
  43. Insights on representational similarity in neural networks with canonical correlation

    Ari S. Morcos, Maithra Raghu, Samy Bengio

    stat.MLcs.AIcs.CVarXiv:1806.05759v32018
  44. Objects that Sound

    Relja Arandjelović, Andrew Zisserman

    cs.CVcs.LGcs.MMarXiv:1712.06651v22017
  45. Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation

    Nai-Xin Zhai, Weihua Cheng, Dexu Yu +9

    cs.CVcs.AIarXiv:2608.21425v12026
  46. Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

    Hangrui Xu, Zhengxian Wu, Yunyao Yu +6

    cs.CVarXiv:2608.21450v12026
  47. End-to-end Learning of Deep Visual Representations for Image Retrieval

    Albert Gordo, Jon Almazan, Jerome Revaud +1

    cs.CVarXiv:1610.07940v22016
  48. Text-Guided Visual Dependency Graph Learning with Cross-Modal Attention Priors

    Fei Wang, Yutong Zhang, Yang Ye +3

    cs.CVarXiv:2608.21443v12026
  49. TransFG: A Transformer Architecture for Fine-grained Recognition

    Ju He, Jie-Neng Chen, Shuai Liu +4

    cs.CVarXiv:2103.07976v52021
  50. FSRNet: End-to-End Learning Face Super-Resolution with Facial Priors

    Yu Chen, Ying Tai, Xiaoming Liu +2

    cs.CVarXiv:1711.10703v12017
  51. The Plan, Not the Decoder: Diagnosing and Repairing Compositional Failure in Reasoning-Augmented Text-to-Image Generation

    Ashritha Gonuguntla

    cs.CVcs.CLcs.LGarXiv:2608.21713v12026
  52. Few-Shot Cross-Dataset Adaptation for Tuberculosis Detection Using DenseNet

    Bidhan Biswas, Shahadat Hossain Sohag, Nabil Ashab +2

    cs.CVarXiv:2608.21427v12026
  53. Implicit Functions in Feature Space for 3D Shape Reconstruction and Completion

    Julian Chibane, Thiemo Alldieck, Gerard Pons-Moll

    cs.CVcs.LGarXiv:2003.01456v22020
  54. Selective Cross-View Consistency for World Action Models: Held-Out Viewpoint Robustness Without Test-Time Camera Information

    Bingqi Huang, Bingchuan Wei, Yingkai Cai +1

    cs.ROcs.CVarXiv:2608.21402v12026
  55. Fader Networks: Manipulating Images by Sliding Attributes

    Guillaume Lample, Neil Zeghidour, Nicolas Usunier +3

    cs.CVarXiv:1706.00409v22017
  56. AI Visual Inspection for Garment Production

    Ray Wai Man Kong, Ding Ning, Theodore Ho Tin Kong

    cs.CVcs.ROarXiv:2608.21426v12026
  57. Boosting Knowledge-based Visual Question Answering with Structured Context Reasoning

    Qiyou Liu, Yong Zhang, Jianjie Luo +2

    cs.CVcs.MMarXiv:2608.21431v12026
  58. EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing

    Yuqian Zhou, Zhenghong Zhou, Zongze Wu +5

    cs.CVcs.GRcs.HCarXiv:2608.21424v12026
  59. ViTexSZ: Heterogeneous Vision-Text Knowledge Distillation for EEG Seizure Detection

    Chenxi Liu, Mingzhao Li, Yicong Liu +4

    cs.CVarXiv:2608.21445v12026
  60. Topology of a Smile: Persistent Homology in Dental Imaging

    Leon Dahlmeier, Sara Kališnik, Albert Mehl +1

    cs.CVmath.ATarXiv:2608.21422v12026