Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

10,501 to 10,560 of 18,866

  1. I$^2$SB: Image-to-Image Schrödinger Bridge

    Guan-Horng Liu, Arash Vahdat, De-An Huang +3

    cs.CVcs.LGstat.MLarXiv:2302.05872v32023
  2. HiFT: Hierarchical Feature Transformer for Aerial Tracking

    Ziang Cao, Changhong Fu, Junjie Ye +2

    cs.CVcs.ROarXiv:2108.00202v32021
  3. Defocus Deblurring Using Dual-Pixel Data

    Abdullah Abuolaim, Michael S. Brown

    eess.IVcs.CVarXiv:2005.00305v32020
  4. Forward Compatible Few-Shot Class-Incremental Learning

    Da-Wei Zhou, Fu-Yun Wang, Han-Jia Ye +3

    cs.CVcs.LGarXiv:2203.06953v12022
  5. CS2-Net: Deep Learning Segmentation of Curvilinear Structures in Medical Imaging

    Lei Mou, Yitian Zhao, Huazhu Fu +9

    eess.IVcs.CVarXiv:2010.07486v22020
  6. Rules of the Road: Predicting Driving Behavior with a Convolutional Model of Semantic Interactions

    Joey Hong, Benjamin Sapp, James Philbin

    cs.CVcs.LGcs.ROarXiv:1906.08945v12019
  7. DeepPET: A deep encoder-decoder network for directly solving the PET reconstruction inverse problem

    Ida Häggström, C. Ross Schmidtlein, Gabriele Campanella +1

    cs.CVphysics.med-pharXiv:1804.07851v22018
  8. DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving

    Erfei Cui, Wenhai Wang, Zhiqi Li +7

    cs.CVarXiv:2312.09245v32023
  9. Neural Face Editing with Intrinsic Image Disentangling

    Zhixin Shu, Ersin Yumer, Sunil Hadap +3

    cs.CVarXiv:1704.04131v12017
  10. Key Points Estimation and Point Instance Segmentation Approach for Lane Detection

    Yeongmin Ko, Younkwan Lee, Shoaib Azam +3

    cs.CVcs.LGeess.IVarXiv:2002.06604v42020
  11. Model Adaptation: Historical Contrastive Learning for Unsupervised Domain Adaptation without Source Data

    Jiaxing Huang, Dayan Guan, Aoran Xiao +1

    cs.CVarXiv:2110.03374v62021
  12. Few-Shot Video Classification via Temporal Alignment

    Kaidi Cao, Jingwei Ji, Zhangjie Cao +2

    cs.CVarXiv:1906.11415v12019
  13. Adversarial Robustness: From Self-Supervised Pre-Training to Fine-Tuning

    Tianlong Chen, Sijia Liu, Shiyu Chang +3

    cs.CVcs.LGarXiv:2003.12862v12020
  14. Modeling the Distribution of Normal Data in Pre-Trained Deep Features for Anomaly Detection

    Oliver Rippel, Patrick Mertens, Dorit Merhof

    cs.CVarXiv:2005.14140v22020
  15. Adversarial Dropout Regularization

    Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada +1

    cs.CVarXiv:1711.01575v32017
  16. Are Multimodal Transformers Robust to Missing Modality?

    Mengmeng Ma, Jian Ren, Long Zhao +2

    cs.CVarXiv:2204.05454v12022
  17. Decoupled Attention Network for Text Recognition

    Tianwei Wang, Yuanzhi Zhu, Lianwen Jin +5

    cs.CVarXiv:1912.10205v12019
  18. RFLA: Gaussian Receptive Field based Label Assignment for Tiny Object Detection

    Chang Xu, Jinwang Wang, Wen Yang +3

    cs.CVarXiv:2208.08738v22022
  19. Attentive Single-Tasking of Multiple Tasks

    Kevis-Kokitsi Maninis, Ilija Radosavovic, Iasonas Kokkinos

    cs.CVarXiv:1904.08918v12019
  20. Weakly Supervised Complementary Parts Models for Fine-Grained Image Classification from the Bottom Up

    Weifeng Ge, Xiangru Lin, Yizhou Yu

    cs.CVarXiv:1903.02827v12019
  21. FaceBoxes: A CPU Real-time Face Detector with High Accuracy

    Shifeng Zhang, Xiangyu Zhu, Zhen Lei +3

    cs.CVarXiv:1708.05234v42017
  22. Instance-Level Salient Object Segmentation

    Guanbin Li, Yuan Xie, Liang Lin +1

    cs.CVarXiv:1704.03604v12017
  23. Beyond Static Features for Temporally Consistent 3D Human Pose and Shape from a Video

    Hongsuk Choi, Gyeongsik Moon, Ju Yong Chang +1

    cs.CVarXiv:2011.08627v42020
  24. DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the Wild

    Xingxun Jiang, Yuan Zong, Wenming Zheng +4

    cs.CVcs.MMarXiv:2008.05924v12020
  25. Shape Prior Deformation for Categorical 6D Object Pose and Size Estimation

    Meng Tian, Marcelo H Ang, Gim Hee Lee

    cs.CVarXiv:2007.08454v12020
  26. An application of cascaded 3D fully convolutional networks for medical image segmentation

    Holger R. Roth, Hirohisa Oda, Xiangrong Zhou +7

    cs.CVarXiv:1803.05431v22018
  27. ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering

    Kan Chen, Jiang Wang, Liang-Chieh Chen +3

    cs.CVarXiv:1511.05960v22015
  28. Editing in Style: Uncovering the Local Semantics of GANs

    Edo Collins, Raja Bala, Bob Price +1

    cs.CVcs.LGarXiv:2004.14367v22020
  29. Perturbed and Strict Mean Teachers for Semi-supervised Semantic Segmentation

    Yuyuan Liu, Yu Tian, Yuanhong Chen +3

    cs.CVarXiv:2111.12903v32021
  30. MetaSDF: Meta-learning Signed Distance Functions

    Vincent Sitzmann, Eric R. Chan, Richard Tucker +2

    cs.CVcs.GRcs.LGarXiv:2006.09662v12020
  31. VideoMatch: Matching based Video Object Segmentation

    Yuan-Ting Hu, Jia-Bin Huang, Alexander G. Schwing

    cs.CVcs.LGarXiv:1809.01123v12018
  32. Accurate Single Stage Detector Using Recurrent Rolling Convolution

    Jimmy Ren, Xiaohao Chen, Jianbo Liu +5

    cs.CVarXiv:1704.05776v12017
  33. SoccerNet: A Scalable Dataset for Action Spotting in Soccer Videos

    Silvio Giancola, Mohieddine Amine, Tarek Dghaily +1

    cs.CVarXiv:1804.04527v22018
  34. Deep Learning on Lie Groups for Skeleton-based Action Recognition

    Zhiwu Huang, Chengde Wan, Thomas Probst +1

    cs.CVarXiv:1612.05877v22016
  35. Self-Supervised Monocular Depth Hints

    Jamie Watson, Michael Firman, Gabriel J. Brostow +1

    cs.CVarXiv:1909.09051v12019
  36. Real-Time MDNet

    Ilchae Jung, Jeany Son, Mooyeol Baek +1

    cs.CVarXiv:1808.08834v12018
  37. Learning to Separate Object Sounds by Watching Unlabeled Video

    Ruohan Gao, Rogerio Feris, Kristen Grauman

    cs.CVcs.MMcs.SDarXiv:1804.01665v22018
  38. Machine Learning-based Lung and Colon Cancer Detection using Deep Feature Extraction and Ensemble Learning

    Md. Alamin Talukder, Md. Manowarul Islam, Md Ashraf Uddin +3

    eess.IVcs.CVcs.LGarXiv:2206.01088v22022
  39. Deep Spectral Clustering using Dual Autoencoder Network

    Xu Yang, Cheng Deng, Feng Zheng +2

    cs.LGcs.CVstat.MLarXiv:1904.13113v12019
  40. Quantifying Radiographic Knee Osteoarthritis Severity using Deep Convolutional Neural Networks

    Joseph Antony, Kevin McGuinness, Noel E O Connor +1

    cs.CVarXiv:1609.02469v12016
  41. GenAttack: Practical Black-box Attacks with Gradient-Free Optimization

    Moustafa Alzantot, Yash Sharma, Supriyo Chakraborty +3

    cs.LGcs.AIcs.CRarXiv:1805.11090v32018
  42. Real-time self-adaptive deep stereo

    Alessio Tonioni, Fabio Tosi, Matteo Poggi +2

    cs.CVarXiv:1810.05424v22018
  43. Navigation World Models

    Amir Bar, Gaoyue Zhou, Danny Tran +2

    cs.CVcs.AIcs.LGarXiv:2412.03572v22024
  44. Robust 3D Hand Pose Estimation in Single Depth Images: from Single-View CNN to Multi-View CNNs

    Liuhao Ge, Hui Liang, Junsong Yuan +1

    cs.CVarXiv:1606.07253v32016
  45. MOSE: A New Dataset for Video Object Segmentation in Complex Scenes

    Henghui Ding, Chang Liu, Shuting He +3

    cs.CVarXiv:2302.01872v12023
  46. Net2Vec: Quantifying and Explaining how Concepts are Encoded by Filters in Deep Neural Networks

    Ruth Fong, Andrea Vedaldi

    cs.CVcs.AIstat.MLarXiv:1801.03454v22018
  47. Whole Slide Images are 2D Point Clouds: Context-Aware Survival Prediction using Patch-based Graph Convolutional Networks

    Richard J. Chen, Ming Y. Lu, Muhammad Shaban +4

    eess.IVcs.CVq-bio.TOarXiv:2107.13048v12021
  48. PixelLM: Pixel Reasoning with Large Multimodal Model

    Zhongwei Ren, Zhicheng Huang, Yunchao Wei +4

    cs.CVarXiv:2312.02228v32023
  49. Balanced Contrastive Learning for Long-Tailed Visual Recognition

    Jianggang Zhu, Zheng Wang, Jingjing Chen +2

    cs.CVcs.LGarXiv:2207.09052v32022
  50. CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

    Seokju Cho, Heeseong Shin, Sunghwan Hong +3

    cs.CVarXiv:2303.11797v22023
  51. Controllable Person Image Synthesis with Attribute-Decomposed GAN

    Yifang Men, Yiming Mao, Yuning Jiang +2

    cs.CVarXiv:2003.12267v42020
  52. REGTR: End-to-end Point Cloud Correspondences with Transformers

    Zi Jian Yew, Gim Hee Lee

    cs.CVarXiv:2203.14517v12022
  53. Advancements in Image Classification using Convolutional Neural Network

    Farhana Sultana, A. Sufian, Paramartha Dutta

    cs.CVcs.AIarXiv:1905.03288v12019
  54. EPSANet: An Efficient Pyramid Squeeze Attention Block on Convolutional Neural Network

    Hu Zhang, Keke Zu, Jian Lu +2

    cs.CVarXiv:2105.14447v22021
  55. Listen to Look: Action Recognition by Previewing Audio

    Ruohan Gao, Tae-Hyun Oh, Kristen Grauman +1

    cs.CVcs.LGcs.SDarXiv:1912.04487v32019
  56. UnrealCV: Connecting Computer Vision to Unreal Engine

    Weichao Qiu, Alan Yuille

    cs.CVarXiv:1609.01326v12016
  57. Wasserstein CNN: Learning Invariant Features for NIR-VIS Face Recognition

    Ran He, Xiang Wu, Zhenan Sun +1

    cs.CVarXiv:1708.02412v12017
  58. VConv-DAE: Deep Volumetric Shape Learning Without Object Labels

    Abhishek Sharma, Oliver Grau, Mario Fritz

    cs.CVcs.GRarXiv:1604.03755v32016
  59. Audio-Driven Emotional Video Portraits

    Xinya Ji, Hang Zhou, Kaisiyuan Wang +4

    cs.CVarXiv:2104.07452v22021
  60. Detecting Curve Text in the Wild: New Dataset and New Solution

    Liu Yuliang, Jin Lianwen, Zhang Shuaitao +1

    cs.CVarXiv:1712.02170v12017