Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
8,521 to 8,580 of 18,867
Learning from Extrinsic and Intrinsic Supervisions for Domain Generalization
Shujun Wang, Lequan Yu, Caizi Li +2
cs.CVarXiv:2007.09316v12020Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System
Penghao Wu, Haiwen Diao, Weichen Fan +3
cs.CVarXiv:2609.01607v12026Multi-scale Domain-adversarial Multiple-instance CNN for Cancer Subtype Classification with Unannotated Histopathological Images
Noriaki Hashimoto, Daisuke Fukushima, Ryoichi Koga +7
cs.CVcs.LGeess.IVarXiv:2001.01599v22020Instance-aware Image and Sentence Matching with Selective Multimodal LSTM
Yan Huang, Wei Wang, Liang Wang
cs.CVarXiv:1611.05588v12016ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
Xionghao Wu, Yijun Yang, Shiyang Zhou +17
cs.CVarXiv:2609.00188v12026LAFITE: Towards Language-Free Training for Text-to-Image Generation
Yufan Zhou, Ruiyi Zhang, Changyou Chen +6
cs.CVcs.LGarXiv:2111.13792v32021CSGNet: Neural Shape Parser for Constructive Solid Geometry
Gopal Sharma, Rishabh Goyal, Difan Liu +2
cs.CVcs.AIarXiv:1712.08290v22017SampleNet: Differentiable Point Cloud Sampling
Itai Lang, Asaf Manor, Shai Avidan
cs.CVarXiv:1912.03663v22019On the Reconstruction of Face Images from Deep Face Templates
Guangcan Mai, Kai Cao, Pong C. Yuen +1
cs.CVarXiv:1703.00832v42017H3-World: Turning Language Understanding into World Control
Danze Chen, Zeqing Wang, Ziyue Lin +2
cs.CVcs.AIarXiv:2609.01560v12026Deep Pixel-wise Binary Supervision for Face Presentation Attack Detection
Anjith George, Sebastien Marcel
cs.CVcs.CRarXiv:1907.04047v12019Cloze Test Helps: Effective Video Anomaly Detection via Learning to Complete Video Events
Guang Yu, Siqi Wang, Zhiping Cai +4
cs.CVcs.LGeess.IVarXiv:2008.11988v12020Document-level Relation Extraction as Semantic Segmentation
Ningyu Zhang, Xiang Chen, Xin Xie +6
cs.CLcs.AIcs.CVarXiv:2106.03618v22021Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
Mingwang Xu, Hui Li, Qingkun Su +6
cs.CVarXiv:2406.08801v22024H-vmunet: High-order Vision Mamba UNet for Medical Image Segmentation
Renkai Wu, Yinghao Liu, Pengchen Liang +1
cs.CVarXiv:2403.13642v12024Vision Language Models in Autonomous Driving: A Survey and Outlook
Xingcheng Zhou, Mingyu Liu, Ekim Yurtsever +4
cs.CVcs.AIarXiv:2310.14414v22023OPUS-V2: Bridging the Gap between Sparse Points and Dense Voxels
Jiabao Wang, Qiang Meng, Liujiang Yan +3
cs.CVarXiv:2608.29187v12026Convolutional Neural Fabrics
Shreyas Saxena, Jakob Verbeek
cs.CVcs.LGcs.NEarXiv:1606.02492v42016Plug and play methods for magnetic resonance imaging (long version)
Rizwan Ahmad, Charles A. Bouman, Gregery T. Buzzard +4
cs.CVarXiv:1903.08616v52019Towards Fully Automated Medical Imaging Code Generation via Validation-based Context Engineering
Zixiao Zhao, Jing Sun, Zhe Hou +5
cs.CVcs.SEarXiv:2608.29016v12026Audio-Visual Scene-Aware Dialog
Huda Alamri, Vincent Cartillier, Abhishek Das +9
cs.CVarXiv:1901.09107v22019RanPAC: Random Projections and Pre-trained Models for Continual Learning
Mark D. McDonnell, Dong Gong, Amin Parveneh +2
cs.LGcs.CVarXiv:2307.02251v32023PointDAN: A Multi-Scale 3D Domain Adaption Network for Point Cloud Representation
Can Qin, Haoxuan You, Lichen Wang +2
cs.CVcs.LGarXiv:1911.02744v12019EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions
Zhiyuan Chen, Jiajiong Cao, Zhiquan Chen +2
cs.CVarXiv:2407.08136v22024Adaptive Multi-Teacher Multi-level Knowledge Distillation
Yuang Liu, Wei Zhang, Jun Wang
cs.CVarXiv:2103.04062v12021Cross-modality image synthesis from unpaired data using CycleGAN: Effects of gradient consistency loss and training data size
Yuta Hiasa, Yoshito Otake, Masaki Takao +5
cs.CVarXiv:1803.06629v32018ReFlowSET: Representation-Aligned Latent Flow Matching for SAR-to-EO Image Translation
Jeonghyeok Do, Seungchul Lee, Munchurl Kim
cs.CVarXiv:2609.00968v12026Cross-domain Contrastive Learning for Unsupervised Domain Adaptation
Rui Wang, Zuxuan Wu, Zejia Weng +3
cs.CVcs.AIcs.LGarXiv:2106.05528v22021CFC-Net: A Critical Feature Capturing Network for Arbitrary-Oriented Object Detection in Remote Sensing Images
Qi Ming, Lingjuan Miao, Zhiqiang Zhou +1
cs.CVarXiv:2101.06849v22021Automatic Radiology Report Generation based on Multi-view Image Fusion and Medical Concept Enrichment
Jianbo Yuan, Haofu Liao, Rui Luo +1
eess.IVcs.CVcs.MMarXiv:1907.09085v22019Learning monocular depth estimation infusing traditional stereo knowledge
Fabio Tosi, Filippo Aleotti, Matteo Poggi +1
cs.CVarXiv:1904.04144v12019Symbiotic Graph Neural Networks for 3D Skeleton-based Human Action Recognition and Motion Prediction
Maosen Li, Siheng Chen, Xu Chen +3
cs.CVarXiv:1910.02212v12019A General Framework for Adversarial Examples with Objectives
Mahmood Sharif, Sruti Bhagavatula, Lujo Bauer +1
cs.CVcs.CRarXiv:1801.00349v22017Memory-Attended Recurrent Network for Video Captioning
Wenjie Pei, Jiyuan Zhang, Xiangrong Wang +3
cs.CVarXiv:1905.03966v12019Segmentation of Bovid Dentition Under Imperfect Annotations: A Comparative Study of Convolutional and Attention Models
Keith G. Mills, Evan B. Sanders, Gregory J. Matthews +1
cs.CVcs.LGarXiv:2608.31052v12026Global Context Vision Transformers
Ali Hatamizadeh, Hongxu Yin, Greg Heinrich +2
cs.CVcs.AIcs.LGarXiv:2206.09959v52022Learn to Match: Automatic Matching Network Design for Visual Tracking
Zhipeng Zhang, Yihao Liu, Xiao Wang +2
cs.CVarXiv:2108.00803v12021Adaptive Sparse Convolutional Networks with Global Context Enhancement for Faster Object Detection on Drone Images
Bowei Du, Yecheng Huang, Jiaxin Chen +1
cs.CVarXiv:2303.14488v12023Marr Revisited: 2D-3D Alignment via Surface Normal Prediction
Aayush Bansal, Bryan Russell, Abhinav Gupta
cs.CVarXiv:1604.01347v12016Stable View Synthesis
Gernot Riegler, Vladlen Koltun
cs.CVarXiv:2011.07233v22020One Adapter, Many Tasks: Task-Conditioned Feature Transformations for Continual Learning
Yunxiang Fu, Meng Lou, Yizhou Yu
cs.CVcs.LGarXiv:2608.31096v12026On the Importance of Noise Scheduling for Diffusion Models
Ting Chen
cs.CVcs.GRcs.LGarXiv:2301.10972v42023Driving on Memory
Christian Löwens, Thorben Funke, Alexandru Paul Condurache
cs.CVcs.LGcs.ROarXiv:2608.31029v12026Zero-Shot Out-of-Distribution Detection Based on the Pre-trained Model CLIP
Sepideh Esmaeilpour, Bing Liu, Eric Robertson +1
cs.CVcs.LGarXiv:2109.02748v32021Findings of the Second Shared Task on Multimodal Machine Translation and Multilingual Image Description
Desmond Elliott, Stella Frank, Loïc Barrault +2
cs.CLcs.CVarXiv:1710.07177v12017UMDFaces: An Annotated Face Dataset for Training Deep Networks
Ankan Bansal, Anirudh Nanduri, Carlos Castillo +2
cs.CVarXiv:1611.01484v22016Orientation-boosted Voxel Nets for 3D Object Recognition
Nima Sedaghat, Mohammadreza Zolfaghari, Ehsan Amiri +1
cs.CVcs.NEarXiv:1604.03351v22016Deep Cross-Modal Audio-Visual Generation
Lele Chen, Sudhanshu Srivastava, Zhiyao Duan +1
cs.CVcs.MMcs.SDarXiv:1704.08292v12017Action Recognition Based on Joint Trajectory Maps with Convolutional Neural Networks
Pichao Wang, Wanqing Li, Chuankun Li +1
cs.CVarXiv:1612.09401v12016Learning JPEG Compression Artifacts for Image Manipulation Detection and Localization
Myung-Joon Kwon, Seung-Hun Nam, In-Jae Yu +2
eess.IVcs.CVcs.LGarXiv:2108.12947v22021Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks
Sihyun Yu, Jihoon Tack, Sangwoo Mo +4
cs.CVcs.LGarXiv:2202.10571v120223D Quasi-Recurrent Neural Network for Hyperspectral Image Denoising
Kaixuan Wei, Ying Fu, Hua Huang
cs.CVarXiv:2003.04547v12020Structured Sparse Subspace Clustering: A Joint Affinity Learning and Subspace Clustering Framework
Chun-Guang Li, Chong You, René Vidal
cs.CVarXiv:1610.05211v22016RAPIQUE: Rapid and Accurate Video Quality Prediction of User Generated Content
Zhengzhong Tu, Xiangxu Yu, Yilin Wang +3
cs.CVcs.MMeess.IVarXiv:2101.10955v22021Retina U-Net: Embarrassingly Simple Exploitation of Segmentation Supervision for Medical Object Detection
Paul F. Jaeger, Simon A. A. Kohl, Sebastian Bickelhaupt +4
cs.CVarXiv:1811.08661v12018STA: Spatial-Temporal Attention for Large-Scale Video-based Person Re-Identification
Yang Fu, Xiaoyang Wang, Yunchao Wei +1
cs.CVarXiv:1811.04129v12018MeteorNet: Deep Learning on Dynamic 3D Point Cloud Sequences
Xingyu Liu, Mengyuan Yan, Jeannette Bohg
cs.CVcs.LGcs.ROarXiv:1910.09165v22019[Extended version] Rethinking Deep Neural Network Ownership Verification: Embedding Passports to Defeat Ambiguity Attacks
Lixin Fan, Kam Woh Ng, Chee Seng Chan
cs.CRcs.CVcs.LGarXiv:1909.07830v32019Video Frame Interpolation Transformer
Zhihao Shi, Xiangyu Xu, Xiaohong Liu +2
cs.CVarXiv:2111.13817v32021Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data
Chaoyi Wu, Xiaoman Zhang, Ya Zhang +2
cs.CVcs.CLarXiv:2308.02463v52023