Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

9,541 to 9,600 of 18,811

  1. All about Structure: Adapting Structural Information across Domains for Boosting Semantic Segmentation

    Wei-Lun Chang, Hui-Po Wang, Wen-Hsiao Peng +1

    cs.CVarXiv:1903.12212v12019
  2. DisturbLabel: Regularizing CNN on the Loss Layer

    Lingxi Xie, Jingdong Wang, Zhen Wei +2

    cs.CVarXiv:1605.00055v12016
  3. Boundary-Guided Camouflaged Object Detection

    Yujia Sun, Shuo Wang, Chenglizhao Chen +1

    cs.CVarXiv:2207.00794v12022
  4. Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation

    Zibo Zhao, Wen Liu, Xin Chen +7

    cs.CVarXiv:2306.17115v22023
  5. Fully Unleashing the Multimodal Attacker: Meta-Adaptive Jailbreaking of Vision-Language Models

    Benlei Cui, Shen Pang, Yuke Wang +7

    cs.CRcs.CVarXiv:2608.27531v12026
  6. Robust Consistent Video Depth Estimation

    Johannes Kopf, Xuejian Rong, Jia-Bin Huang

    cs.CVarXiv:2012.05901v22020
  7. Distilling Object Detectors via Decoupled Features

    Jianyuan Guo, Kai Han, Yunhe Wang +4

    cs.CVarXiv:2103.14475v12021
  8. Relational Embedding for Few-Shot Classification

    Dahyun Kang, Heeseung Kwon, Juhong Min +1

    cs.CVarXiv:2108.09666v12021
  9. Large-scale interactive object segmentation with human annotators

    Rodrigo Benenson, Stefan Popov, Vittorio Ferrari

    cs.CVarXiv:1903.10830v22019
  10. HOME: Heatmap Output for future Motion Estimation

    Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou +2

    cs.CVcs.ROarXiv:2105.10968v22021
  11. CIFAKE: Image Classification and Explainable Identification of AI-Generated Synthetic Images

    Jordan J. Bird, Ahmad Lotfi

    cs.CVcs.AIcs.LGarXiv:2303.14126v12023
  12. LCM-LoRA: A Universal Stable-Diffusion Acceleration Module

    Simian Luo, Yiqin Tan, Suraj Patil +6

    cs.CVcs.LGarXiv:2311.05556v12023
  13. Referring Transformer: A One-step Approach to Multi-task Visual Grounding

    Muchen Li, Leonid Sigal

    cs.CVarXiv:2106.03089v22021
  14. Graphical Contrastive Losses for Scene Graph Parsing

    Ji Zhang, Kevin J. Shih, Ahmed Elgammal +2

    cs.CVarXiv:1903.02728v52019
  15. Attention-Driven Dynamic Graph Convolutional Network for Multi-Label Image Recognition

    Jin Ye, Junjun He, Xiaojiang Peng +2

    cs.CVarXiv:2012.02994v12020
  16. Generating 3D structures from a 2D slice with GAN-based dimensionality expansion

    Steve Kench, Samuel J. Cooper

    cs.CVcs.LGarXiv:2102.07708v12021
  17. Distribution Matching for Heterogeneous Multi-Task Learning: a Large-scale Face Study

    Dimitrios Kollias, Viktoriia Sharmanska, Stefanos Zafeiriou

    cs.CVarXiv:2105.03790v12021
  18. VidParse: Online Parsing of Egocentric Procedures Like a Pro

    Anubhav Gupta, Archit Kambhamettu, Vatsal Agarwal +2

    cs.CVarXiv:2608.27562v12026
  19. Who2com: Collaborative Perception via Learnable Handshake Communication

    Yen-Cheng Liu, Junjiao Tian, Chih-Yao Ma +3

    cs.CVcs.MAcs.ROarXiv:2003.09575v12020
  20. ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding

    Le Xue, Ning Yu, Shu Zhang +8

    cs.CVarXiv:2305.08275v42023
  21. DoveNet: Deep Image Harmonization via Domain Verification

    Wenyan Cong, Jianfu Zhang, Li Niu +4

    cs.CVarXiv:1911.13239v32019
  22. Multiscale Point Cloud Geometry Compression

    Jianqiang Wang, Dandan Ding, Zhu Li +1

    eess.IVcs.CVarXiv:2011.03799v12020
  23. A Deeper Analysis of Block-Sparse Featurizers

    Alexandru-Iulius Jerpelea, Amith Ananthram

    cs.LGcs.CVarXiv:2608.27515v12026
  24. Learned Video Compression

    Oren Rippel, Sanjay Nair, Carissa Lew +3

    eess.IVcs.CVcs.LGarXiv:1811.06981v12018
  25. GRF: Learning a General Radiance Field for 3D Representation and Rendering

    Alex Trevithick, Bo Yang

    cs.CVcs.AIcs.GRarXiv:2010.04595v32020
  26. Deep Networks with Internal Selective Attention through Feedback Connections

    Marijn Stollenga, Jonathan Masci, Faustino Gomez +1

    cs.CVcs.LGcs.NEarXiv:1407.3068v22014
  27. Brain Tumor Segmentation using an Ensemble of 3D U-Nets and Overall Survival Prediction using Radiomic Features

    Xue Feng, Nicholas Tustison, Craig Meyer

    cs.CVarXiv:1812.01049v12018
  28. Analysing Affective Behavior in the second ABAW2 Competition

    Dimitrios Kollias, Irene Kotsia, Elnar Hajiyev +1

    cs.CVarXiv:2106.15318v22021
  29. Learning Residual Images for Face Attribute Manipulation

    Wei Shen, Rujie Liu

    cs.CVarXiv:1612.05363v22016
  30. Recurrent Autoregressive Networks for Online Multi-Object Tracking

    Kuan Fang, Yu Xiang, Xiaocheng Li +1

    cs.CVcs.AIcs.LGarXiv:1711.02741v22017
  31. Multiscale Deep Equilibrium Models

    Shaojie Bai, Vladlen Koltun, J. Zico Kolter

    cs.LGcs.CVstat.MLarXiv:2006.08656v22020
  32. Where am I looking at? Joint Location and Orientation Estimation by Cross-View Matching

    Yujiao Shi, Xin Yu, Dylan Campbell +1

    cs.CVarXiv:2005.03860v12020
  33. Fast Symmetric Diffeomorphic Image Registration with Convolutional Neural Networks

    Tony C. W. Mok, Albert C. S. Chung

    cs.CVarXiv:2003.09514v32020
  34. MGFN: Magnitude-Contrastive Glance-and-Focus Network for Weakly-Supervised Video Anomaly Detection

    Yingxian Chen, Zhengzhe Liu, Baoheng Zhang +3

    cs.CVarXiv:2211.15098v12022
  35. Re-Imagen: Retrieval-Augmented Text-to-Image Generator

    Wenhu Chen, Hexiang Hu, Chitwan Saharia +1

    cs.CVcs.AIcs.LGarXiv:2209.14491v32022
  36. Neural-PIL: Neural Pre-Integrated Lighting for Reflectance Decomposition

    Mark Boss, Varun Jampani, Raphael Braun +3

    cs.CVcs.GRcs.LGarXiv:2110.14373v12021
  37. ESIR: End-to-end Scene Text Recognition via Iterative Image Rectification

    Fangneng Zhan, Shijian Lu

    cs.CVarXiv:1812.05824v32018
  38. Appearance-based Gaze Estimation With Deep Learning: A Review and Benchmark

    Yihua Cheng, Haofei Wang, Yiwei Bao +1

    cs.CVarXiv:2104.12668v22021
  39. Unsupervised Learning of Shape and Pose with Differentiable Point Clouds

    Eldar Insafutdinov, Alexey Dosovitskiy

    cs.CVcs.LGarXiv:1810.09381v12018
  40. CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model

    Zhengyi Wang, Yikai Wang, Yifei Chen +6

    cs.CVcs.LGarXiv:2403.05034v12024
  41. VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

    Yue Fan, Xiaojian Ma, Rujie Wu +4

    cs.CVarXiv:2403.11481v22024
  42. A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection

    Zhoupeng Guo, Xinjie Yao, Yunqi Zhu +6

    cs.CVcs.MMarXiv:2608.27997v12026
  43. An Analysis of Visual Question Answering Algorithms

    Kushal Kafle, Christopher Kanan

    cs.CVcs.AIcs.CLarXiv:1703.09684v22017
  44. 3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting

    Zhiyin Qian, Shaofei Wang, Marko Mihajlovic +2

    cs.CVarXiv:2312.09228v32023
  45. Hierarchically Decomposed Graph Convolutional Networks for Skeleton-Based Action Recognition

    Jungho Lee, Minhyeok Lee, Dogyoon Lee +1

    cs.CVarXiv:2208.10741v32022
  46. Deep TEN: Texture Encoding Network

    Hang Zhang, Jia Xue, Kristin Dana

    cs.CVarXiv:1612.02844v12016
  47. OpenGait: Revisiting Gait Recognition Toward Better Practicality

    Chao Fan, Junhao Liang, Chuanfu Shen +3

    cs.CVarXiv:2211.06597v32022
  48. Mask-ShadowGAN: Learning to Remove Shadows from Unpaired Data

    Xiaowei Hu, Yitong Jiang, Chi-Wing Fu +1

    cs.CVcs.MMarXiv:1903.10683v32019
  49. BigHand2.2M Benchmark: Hand Pose Dataset and State of the Art Analysis

    Shanxin Yuan, Qi Ye, Bjorn Stenger +2

    cs.CVarXiv:1704.02612v22017
  50. Self-Ensembling with GAN-based Data Augmentation for Domain Adaptation in Semantic Segmentation

    Jaehoon Choi, Taekyung Kim, Changick Kim

    cs.CVarXiv:1909.00589v12019
  51. Unsupervised Domain Adaptation via Structured Prediction Based Selective Pseudo-Labeling

    Qian Wang, Toby P. Breckon

    cs.LGcs.CVcs.MMarXiv:1911.07982v12019
  52. Video Inpainting of Complex Scenes

    Alasdair Newson, Andrés Almansa, Matthieu Fradet +2

    cs.CVcs.MMeess.IVarXiv:1503.05528v22015
  53. Clustering with Deep Learning: Taxonomy and New Methods

    Elie Aljalbout, Vladimir Golkov, Yawar Siddiqui +2

    cs.LGcs.AIcs.CVarXiv:1801.07648v22018
  54. CVAE-GAN: Fine-Grained Image Generation through Asymmetric Training

    Jianmin Bao, Dong Chen, Fang Wen +2

    cs.CVarXiv:1703.10155v22017
  55. shapeDTW: shape Dynamic Time Warping

    Jiaping Zhao, Laurent Itti

    cs.CVarXiv:1606.01601v12016
  56. A Fine-Grained Analysis on Distribution Shift

    Olivia Wiles, Sven Gowal, Florian Stimberg +4

    cs.LGcs.CVarXiv:2110.11328v22021
  57. Incremental Few-Shot Object Detection

    Juan-Manuel Perez-Rua, Xiatian Zhu, Timothy Hospedales +1

    cs.CVarXiv:2003.04668v22020
  58. Finding Covid-19 from Chest X-rays using Deep Learning on a Small Dataset

    Lawrence O. Hall, Rahul Paul, Dmitry B. Goldgof +1

    eess.IVcs.CVcs.LGarXiv:2004.02060v42020
  59. MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound

    Rowan Zellers, Jiasen Lu, Ximing Lu +7

    cs.CVcs.CLcs.LGarXiv:2201.02639v42022
  60. Discriminatively Boosted Image Clustering with Fully Convolutional Auto-Encoders

    Fengfu Li, Hong Qiao, Bo Zhang +1

    cs.CVcs.LGarXiv:1703.07980v12017