Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,061 to 3,120 of 18,855

  1. Context-aware stacked convolutional neural networks for classification of breast carcinomas in whole-slide histopathology images

    Babak Ehteshami Bejnordi, Guido Zuidhof, Maschenka Balkenhol +6

    cs.CVarXiv:1705.03678v12017
  2. Best sources forward: domain generalization through source-specific nets

    Massimiliano Mancini, Samuel Rota Bulò, Barbara Caputo +1

    cs.CVcs.LGstat.MLarXiv:1806.05810v12018
  3. ADMM-NN: An Algorithm-Hardware Co-Design Framework of DNNs Using Alternating Direction Method of Multipliers

    Ao Ren, Tianyun Zhang, Shaokai Ye +5

    cs.LGcs.AIcs.ARarXiv:1812.11677v12018
  4. Learning Deep Multi-Level Similarity for Thermal Infrared Object Tracking

    Qiao Liu, Xin Li, Zhenyu He +3

    cs.CVarXiv:1906.03568v12019
  5. Episode-based Prototype Generating Network for Zero-Shot Learning

    Yunlong Yu, Zhong Ji, Zhongfei Zhang +1

    cs.CVarXiv:1909.03360v22019
  6. Photorealistic Monocular 3D Reconstruction of Humans Wearing Clothing

    Thiemo Alldieck, Mihai Zanfir, Cristian Sminchisescu

    cs.CVarXiv:2204.08906v12022
  7. Viral Pneumonia Screening on Chest X-ray Images Using Confidence-Aware Anomaly Detection

    Jianpeng Zhang, Yutong Xie, Guansong Pang +8

    eess.IVcs.CVarXiv:2003.12338v42020
  8. Robust, Deep and Inductive Anomaly Detection

    Raghavendra Chalapathy, Aditya Krishna Menon, Sanjay Chawla

    cs.LGcs.CVstat.MLarXiv:1704.06743v32017
  9. HyperMorph: Amortized Hyperparameter Learning for Image Registration

    Andrew Hoopes, Malte Hoffmann, Bruce Fischl +2

    cs.CVeess.IVarXiv:2101.01035v22021
  10. Black-box Adversarial Attacks on Video Recognition Models

    Linxi Jiang, Xingjun Ma, Shaoxiang Chen +2

    cs.LGcs.CRcs.CVarXiv:1904.05181v22019
  11. Representation Similarity Analysis for Efficient Task taxonomy & Transfer Learning

    Kshitij Dwivedi, Gemma Roig

    cs.CVcs.AIcs.LGarXiv:1904.11740v12019
  12. Improving Landmark Localization with Semi-Supervised Learning

    Sina Honari, Pavlo Molchanov, Stephen Tyree +3

    cs.CVarXiv:1709.01591v72017
  13. Zoom Better to See Clearer: Human and Object Parsing with Hierarchical Auto-Zoom Net

    Fangting Xia, Peng Wang, Liang-Chieh Chen +1

    cs.CVcs.LGarXiv:1511.06881v52015
  14. Spatial Memory for Context Reasoning in Object Detection

    Xinlei Chen, Abhinav Gupta

    cs.CVarXiv:1704.04224v12017
  15. Dual Illumination Estimation for Robust Exposure Correction

    Qing Zhang, Yongwei Nie, Wei-Shi Zheng

    cs.CVcs.GRarXiv:1910.13688v12019
  16. Spatio-Temporal Covariance Descriptors for Action and Gesture Recognition

    Andres Sanin, Conrad Sanderson, Mehrtash T. Harandi +1

    cs.CVcs.HCarXiv:1303.6021v12013
  17. Rich Human Feedback for Text-to-Image Generation

    Youwei Liang, Junfeng He, Gang Li +15

    cs.CVarXiv:2312.10240v22023
  18. Semantic Instance Annotation of Street Scenes by 3D to 2D Label Transfer

    Jun Xie, Martin Kiefel, Ming-Ting Sun +1

    cs.CVarXiv:1511.03240v22015
  19. Unsupervised Multi-Modal Image Registration via Geometry Preserving Image-to-Image Translation

    Moab Arar, Yiftach Ginger, Dov Danon +3

    cs.CVarXiv:2003.08073v12020
  20. Interventional Video Grounding with Dual Contrastive Learning

    Guoshun Nan, Rui Qiao, Yao Xiao +4

    cs.CVcs.CLarXiv:2106.11013v22021
  21. ResDiff: Combining CNN and Diffusion Model for Image Super-Resolution

    Shuyao Shang, Zhengyang Shan, Guangxing Liu +4

    cs.CVarXiv:2303.08714v32023
  22. A Short Note on the Kinetics-700-2020 Human Action Dataset

    Lucas Smaira, João Carreira, Eric Noland +3

    cs.CVcs.LGarXiv:2010.10864v12020
  23. Large-Scale Visual Speech Recognition

    Brendan Shillingford, Yannis Assael, Matthew W. Hoffman +12

    cs.CVcs.LGarXiv:1807.05162v32018
  24. Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation

    William Shen, Ge Yang, Alan Yu +3

    cs.CVcs.AIcs.CLarXiv:2308.07931v22023
  25. Heterogeneous Domain Generalization via Domain Mixup

    Yufei Wang, Haoliang Li, Alex C. Kot

    cs.CVarXiv:2009.05448v12020
  26. SlotFormer: Unsupervised Visual Dynamics Simulation with Object-Centric Models

    Ziyi Wu, Nikita Dvornik, Klaus Greff +2

    cs.CVcs.AIcs.LGarXiv:2210.05861v22022
  27. Generating 3D People in Scenes without People

    Yan Zhang, Mohamed Hassan, Heiko Neumann +2

    cs.CVarXiv:1912.02923v32019
  28. Invertible Image Signal Processing

    Yazhou Xing, Zian Qian, Qifeng Chen

    eess.IVcs.CVarXiv:2103.15061v22021
  29. Group-aware Contrastive Regression for Action Quality Assessment

    Xumin Yu, Yongming Rao, Wenliang Zhao +2

    cs.CVcs.AIcs.LGarXiv:2108.07797v12021
  30. SUES-200: A Multi-height Multi-scene Cross-view Image Benchmark Across Drone and Satellite

    Runzhe Zhu, Ling Yin, Mingze Yang +3

    cs.CVeess.IVarXiv:2204.10704v22022
  31. What to Hide from Your Students: Attention-Guided Masked Image Modeling

    Ioannis Kakogeorgiou, Spyros Gidaris, Bill Psomas +4

    cs.CVarXiv:2203.12719v22022
  32. PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle Adjustment

    Jianyuan Wang, Christian Rupprecht, David Novotny

    cs.CVarXiv:2306.15667v42023
  33. Deep Multiview Clustering by Contrasting Cluster Assignments

    Jie Chen, Hua Mao, Wai Lok Woo +1

    cs.CVarXiv:2304.10769v42023
  34. OpenPifPaf: Composite Fields for Semantic Keypoint Detection and Spatio-Temporal Association

    Sven Kreiss, Lorenzo Bertoni, Alexandre Alahi

    cs.CVarXiv:2103.02440v22021
  35. Objects2action: Classifying and localizing actions without any video example

    Mihir Jain, Jan C. van Gemert, Thomas Mensink +1

    cs.CVarXiv:1510.06939v12015
  36. Partial Order Pruning: for Best Speed/Accuracy Trade-off in Neural Architecture Search

    Xin Li, Yiming Zhou, Zheng Pan +1

    cs.CVarXiv:1903.03777v22019
  37. Physion: Evaluating Physical Prediction from Vision in Humans and Machines

    Daniel M. Bear, Elias Wang, Damian Mrowca +12

    cs.AIcs.CVarXiv:2106.08261v32021
  38. Domain Adaptive Semantic Segmentation with Self-Supervised Depth Estimation

    Qin Wang, Dengxin Dai, Lukas Hoyer +2

    cs.CVarXiv:2104.13613v22021
  39. An Image Patch is a Wave: Phase-Aware Vision MLP

    Yehui Tang, Kai Han, Jianyuan Guo +4

    cs.CVarXiv:2111.12294v52021
  40. Video Instance Segmentation using Inter-Frame Communication Transformers

    Sukjun Hwang, Miran Heo, Seoung Wug Oh +1

    cs.CVarXiv:2106.03299v12021
  41. Counting Everyday Objects in Everyday Scenes

    Prithvijit Chattopadhyay, Ramakrishna Vedantam, Ramprasaath R. Selvaraju +2

    cs.CVarXiv:1604.03505v32016
  42. DragAnything: Motion Control for Anything using Entity Representation

    Weijia Wu, Zhuang Li, Yuchao Gu +7

    cs.CVarXiv:2403.07420v32024
  43. Graph Contrastive Clustering

    Huasong Zhong, Jianlong Wu, Chong Chen +5

    cs.CVarXiv:2104.01429v12021
  44. RAFT-3D: Scene Flow using Rigid-Motion Embeddings

    Zachary Teed, Jia Deng

    cs.CVarXiv:2012.00726v22020
  45. Neural Design Network: Graphic Layout Generation with Constraints

    Hsin-Ying Lee, Lu Jiang, Irfan Essa +4

    cs.CVarXiv:1912.09421v22019
  46. Hetero-Modal Variational Encoder-Decoder for Joint Modality Completion and Segmentation

    Reuben Dorent, Samuel Joutard, Marc Modat +2

    eess.IVcs.CVarXiv:1907.11150v12019
  47. Bridging Composite and Real: Towards End-to-end Deep Image Matting

    Jizhizi Li, Jing Zhang, Stephen J. Maybank +1

    cs.CVcs.LGeess.IVarXiv:2010.16188v32020
  48. Spatially Conditioned Graphs for Detecting Human-Object Interactions

    Frederic Z. Zhang, Dylan Campbell, Stephen Gould

    cs.CVcs.AIcs.LGarXiv:2012.06060v32020
  49. Panoptic Scene Graph Generation

    Jingkang Yang, Yi Zhe Ang, Zujin Guo +3

    cs.CVcs.AIcs.CLarXiv:2207.11247v12022
  50. Fast Exact Search in Hamming Space with Multi-Index Hashing

    Mohammad Norouzi, Ali Punjani, David J. Fleet

    cs.CVcs.AIcs.DSarXiv:1307.2982v32013
  51. Content Adaptive and Error Propagation Aware Deep Video Compression

    Guo Lu, Chunlei Cai, Xiaoyun Zhang +4

    eess.IVcs.CVarXiv:2003.11282v12020
  52. FPConv: Learning Local Flattening for Point Convolution

    Yiqun Lin, Zizheng Yan, Haibin Huang +4

    cs.CVarXiv:2002.10701v32020
  53. Decoupling Human and Camera Motion from Videos in the Wild

    Vickie Ye, Georgios Pavlakos, Jitendra Malik +1

    cs.CVarXiv:2302.12827v22023
  54. Time Lens++: Event-based Frame Interpolation with Parametric Non-linear Flow and Multi-scale Fusion

    Stepan Tulyakov, Alfredo Bochicchio, Daniel Gehrig +3

    cs.CVarXiv:2203.17191v22022
  55. Learning Data-driven Reflectance Priors for Intrinsic Image Decomposition

    Tinghui Zhou, Philipp Krähenbühl, Alexei A. Efros

    cs.CVarXiv:1510.02413v12015
  56. 3D-LMNet: Latent Embedding Matching for Accurate and Diverse 3D Point Cloud Reconstruction from a Single Image

    Priyanka Mandikal, K L Navaneet, Mayank Agarwal +1

    cs.CVarXiv:1807.07796v22018
  57. Learning Object Relation Graph and Tentative Policy for Visual Navigation

    Heming Du, Xin Yu, Liang Zheng

    cs.CVarXiv:2007.11018v12020
  58. Interventional Bag Multi-Instance Learning On Whole-Slide Pathological Images

    Tiancheng Lin, Zhimiao Yu, Hongyu Hu +2

    cs.CVarXiv:2303.06873v12023
  59. Deep Metric Learning with BIER: Boosting Independent Embeddings Robustly

    Michael Opitz, Georg Waltner, Horst Possegger +1

    cs.CVarXiv:1801.04815v12018
  60. Fashion Landmark Detection in the Wild

    Ziwei Liu, Sijie Yan, Ping Luo +2

    cs.CVarXiv:1608.03049v12016