Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

10,021 to 10,080 of 18,867

  1. 3D UX-Net: A Large Kernel Volumetric ConvNet Modernizing Hierarchical Transformer for Medical Image Segmentation

    Ho Hin Lee, Shunxing Bao, Yuankai Huo +1

    cs.CVcs.LGarXiv:2209.15076v42022
  2. Deep Kinematic Pose Regression

    Xingyi Zhou, Xiao Sun, Wei Zhang +2

    cs.CVarXiv:1609.05317v12016
  3. Tracking Anything with Decoupled Video Segmentation

    Ho Kei Cheng, Seoung Wug Oh, Brian Price +2

    cs.CVarXiv:2309.03903v12023
  4. Frequency Separation for Real-World Super-Resolution

    Manuel Fritsche, Shuhang Gu, Radu Timofte

    eess.IVcs.CVarXiv:1911.07850v12019
  5. An Enhanced Deep Feature Representation for Person Re-identification

    Shangxuan Wu, Ying-Cong Chen, Xiang Li +3

    cs.CVarXiv:1604.07807v22016
  6. Grid-GCN for Fast and Scalable Point Cloud Learning

    Qiangeng Xu, Xudong Sun, Cho-Ying Wu +2

    cs.CVcs.LGarXiv:1912.02984v52019
  7. SNARF: Differentiable Forward Skinning for Animating Non-Rigid Neural Implicit Shapes

    Xu Chen, Yufeng Zheng, Michael J. Black +2

    cs.CVarXiv:2104.03953v32021
  8. Zoom-in-Net: Deep Mining Lesions for Diabetic Retinopathy Detection

    Zhe Wang, Yanxin Yin, Jianping Shi +3

    cs.CVarXiv:1706.04372v12017
  9. Multi-Scale Geometric Consistency Guided Multi-View Stereo

    Qingshan Xu, Wenbing Tao

    cs.CVarXiv:1904.08103v12019
  10. Uncertainty Estimates and Multi-Hypotheses Networks for Optical Flow

    Eddy Ilg, Özgün Çiçek, Silvio Galesso +4

    cs.CVarXiv:1802.07095v42018
  11. You said that?

    Joon Son Chung, Amir Jamaludin, Andrew Zisserman

    cs.CVarXiv:1705.02966v22017
  12. Neural Head Avatars from Monocular RGB Videos

    Philip-William Grassal, Malte Prinzler, Titus Leistner +3

    cs.CVcs.GRarXiv:2112.01554v22021
  13. Overcoming Language Priors in Visual Question Answering with Adversarial Regularization

    Sainandan Ramakrishnan, Aishwarya Agrawal, Stefan Lee

    cs.CVarXiv:1810.03649v22018
  14. DeepGMR: Learning Latent Gaussian Mixture Models for Registration

    Wentao Yuan, Ben Eckart, Kihwan Kim +3

    cs.CVarXiv:2008.09088v12020
  15. Soft Threshold Weight Reparameterization for Learnable Sparsity

    Aditya Kusupati, Vivek Ramanujan, Raghav Somani +4

    cs.LGcs.CVstat.MLarXiv:2002.03231v92020
  16. From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation

    Rit Gangopadhyay, Alex Wong

    cs.CVcs.AIarXiv:2608.27860v12026
  17. MinkLoc3D: Point Cloud Based Large-Scale Place Recognition

    Jacek Komorowski

    cs.CVarXiv:2011.04530v12020
  18. DexGraspNet: A Large-Scale Robotic Dexterous Grasp Dataset for General Objects Based on Simulation

    Ruicheng Wang, Jialiang Zhang, Jiayi Chen +4

    cs.ROcs.CVarXiv:2210.02697v22022
  19. Sequential Attend, Infer, Repeat: Generative Modelling of Moving Objects

    Adam R. Kosiorek, Hyunjik Kim, Ingmar Posner +1

    cs.LGcs.CVstat.MLarXiv:1806.01794v22018
  20. Semi-Supervised Semantic Segmentation with Pixel-Level Contrastive Learning from a Class-wise Memory Bank

    Inigo Alonso, Alberto Sabater, David Ferstl +2

    cs.CVarXiv:2104.13415v32021
  21. EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

    Linrui Tian, Qi Wang, Bang Zhang +1

    cs.CVarXiv:2402.17485v32024
  22. Pose-guided Feature Disentangling for Occluded Person Re-identification Based on Transformer

    Tao Wang, Hong Liu, Pinhao Song +2

    cs.CVarXiv:2112.02466v22021
  23. DeepLPF: Deep Local Parametric Filters for Image Enhancement

    Sean Moran, Pierre Marza, Steven McDonagh +2

    cs.CVarXiv:2003.13985v12020
  24. MoFusion: A Framework for Denoising-Diffusion-based Motion Synthesis

    Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik +1

    cs.CVarXiv:2212.04495v32022
  25. Deep Learning for Unsupervised Anomaly Localization in Industrial Images: A Survey

    Xian Tao, Xinyi Gong, Xin Zhang +2

    cs.CVarXiv:2207.10298v12022
  26. MultiPoseNet: Fast Multi-Person Pose Estimation using Pose Residual Network

    Muhammed Kocabas, Salih Karagoz, Emre Akbas

    cs.CVarXiv:1807.04067v12018
  27. CARDINAL Predicts Cardiovascular Risk From Non-contrast Cardiac CT

    Roy Gabriel, Nattakorn Kittisut, Jamshid Hassanpour +6

    eess.IVcs.AIcs.CVarXiv:2608.27690v12026
  28. Self-supervised Multi-level Face Model Learning for Monocular Reconstruction at over 250 Hz

    Ayush Tewari, Michael Zollhöfer, Pablo Garrido +4

    cs.CVarXiv:1712.02859v22017
  29. VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

    Ziyang Wang, Shoubin Yu, Elias Stengel-Eskin +4

    cs.CVcs.AIcs.CLarXiv:2405.19209v32024
  30. Do Deep Neural Networks Learn Facial Action Units When Doing Expression Recognition?

    Pooya Khorrami, Tom Le Paine, Thomas S. Huang

    cs.CVcs.LGcs.NEarXiv:1510.02969v32015
  31. Railroad is not a Train: Saliency as Pseudo-pixel Supervision for Weakly Supervised Semantic Segmentation

    Seungho Lee, Minhyun Lee, Jongwuk Lee +1

    cs.CVarXiv:2105.08965v12021
  32. USE-Net: incorporating Squeeze-and-Excitation blocks into U-Net for prostate zonal segmentation of multi-institutional MRI datasets

    Leonardo Rundo, Changhee Han, Yudai Nagano +12

    cs.CVcs.LGarXiv:1904.08254v22019
  33. FVeinSyn: Synthetic Finger Vein Image Generator

    Yifan Wang, Jie Gui, Adams Wai Kin Kong +6

    cs.CVcs.AIarXiv:2608.27527v12026
  34. Segment and Track Anything

    Yangming Cheng, Liulei Li, Yuanyou Xu +4

    cs.CVarXiv:2305.06558v12023
  35. Taming 3DGS: High-Quality Radiance Fields with Limited Resources

    Saswat Subhajyoti Mallick, Rahul Goel, Bernhard Kerbl +3

    cs.CVcs.GRarXiv:2406.15643v12024
  36. Training Neural Networks with Local Error Signals

    Arild Nøkland, Lars Hiller Eidnes

    stat.MLcs.CVcs.LGarXiv:1901.06656v22019
  37. TI-POOLING: transformation-invariant pooling for feature learning in Convolutional Neural Networks

    Dmitry Laptev, Nikolay Savinov, Joachim M. Buhmann +1

    cs.CVarXiv:1604.06318v22016
  38. Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification

    Alexandre L. M. Levada

    cs.LGcs.AIcs.CVarXiv:2608.27634v12026
  39. Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP

    Qihang Yu, Ju He, Xueqing Deng +2

    cs.CVarXiv:2308.02487v22023
  40. T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

    Kaiyi Huang, Chengqi Duan, Kaiyue Sun +3

    cs.CVarXiv:2307.06350v32023
  41. Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge

    Md Monjurul Ahsan Prodhan, Md Nour Hossain

    cs.CVcs.AIcs.LGarXiv:2608.27633v12026
  42. Segment Anything Is Not Always Perfect: An Investigation of SAM on Different Real-world Applications

    Wei Ji, Jingjing Li, Qi Bi +3

    cs.CVarXiv:2304.05750v32023
  43. Quanta Perception as Probabilistic Events

    Varun Sundar, Pavan Thodima, Sacha Jungerman +1

    cs.CVcs.AIarXiv:2608.27584v12026
  44. Woodpecker: Hallucination Correction for Multimodal Large Language Models

    Shukang Yin, Chaoyou Fu, Sirui Zhao +7

    cs.CVcs.AIcs.CLarXiv:2310.16045v22023
  45. On Translation Invariance in CNNs: Convolutional Layers can Exploit Absolute Spatial Location

    Osman Semih Kayhan, Jan C. van Gemert

    cs.CVcs.LGeess.IVarXiv:2003.07064v22020
  46. Rice Diseases Detection and Classification Using Attention Based Neural Network and Bayesian Optimization

    Yibin Wang, Haifeng Wang, Zhaohua Peng

    cs.CVarXiv:2201.00893v12022
  47. SynthMorph: learning contrast-invariant registration without acquired images

    Malte Hoffmann, Benjamin Billot, Douglas N. Greve +3

    eess.IVcs.CVq-bio.NCarXiv:2004.10282v42020
  48. MAGE: MAsked Generative Encoder to Unify Representation Learning and Image Synthesis

    Tianhong Li, Huiwen Chang, Shlok Kumar Mishra +3

    cs.CVarXiv:2211.09117v22022
  49. Diffusion Models Are Real-Time Game Engines

    Dani Valevski, Yaniv Leviathan, Moab Arar +1

    cs.LGcs.AIcs.CVarXiv:2408.14837v22024
  50. Neighbourhood Watch: Referring Expression Comprehension via Language-guided Graph Attention Networks

    Peng Wang, Qi Wu, Jiewei Cao +3

    cs.CVarXiv:1812.04794v12018
  51. PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world Learning

    Xiangyang Zhu, Renrui Zhang, Bowei He +5

    cs.CVarXiv:2211.11682v22022
  52. PolyNet: A Pursuit of Structural Diversity in Very Deep Networks

    Xingcheng Zhang, Zhizhong Li, Chen Change Loy +1

    cs.CVarXiv:1611.05725v22016
  53. EGE-UNet: an Efficient Group Enhanced UNet for skin lesion segmentation

    Jiacheng Ruan, Mingye Xie, Jingsheng Gao +2

    eess.IVcs.CVarXiv:2307.08473v12023
  54. Semantic Image Synthesis via Adversarial Learning

    Hao Dong, Simiao Yu, Chao Wu +1

    cs.CVarXiv:1707.06873v12017
  55. Unlearnable Examples: Making Personal Data Unexploitable

    Hanxun Huang, Xingjun Ma, Sarah Monazam Erfani +2

    cs.LGcs.CRcs.CVarXiv:2101.04898v22021
  56. Video Object Segmentation with Language Referring Expressions

    Anna Khoreva, Anna Rohrbach, Bernt Schiele

    cs.CVarXiv:1803.08006v32018
  57. TruFor: Leveraging all-round clues for trustworthy image forgery detection and localization

    Fabrizio Guillaro, Davide Cozzolino, Avneesh Sud +2

    cs.CVarXiv:2212.10957v32022
  58. ArtTrack: Articulated Multi-person Tracking in the Wild

    Eldar Insafutdinov, Mykhaylo Andriluka, Leonid Pishchulin +4

    cs.CVarXiv:1612.01465v32016
  59. Destroy Me: Automatic Artifact Generation for Histopathology Images

    Zuzanna Krawczyk-Borysiak, Adam Krawczyk, Mateusz Miller +4

    eess.IVcs.AIcs.CVarXiv:2608.27516v12026
  60. SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition

    Zhi Qiao, Yu Zhou, Dongbao Yang +2

    cs.CVarXiv:2005.10977v12020