Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

8,521 to 8,580 of 18,867

  1. Learning from Extrinsic and Intrinsic Supervisions for Domain Generalization

    Shujun Wang, Lequan Yu, Caizi Li +2

    cs.CVarXiv:2007.09316v12020
  2. Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System

    Penghao Wu, Haiwen Diao, Weichen Fan +3

    cs.CVarXiv:2609.01607v12026
  3. Multi-scale Domain-adversarial Multiple-instance CNN for Cancer Subtype Classification with Unannotated Histopathological Images

    Noriaki Hashimoto, Daisuke Fukushima, Ryoichi Koga +7

    cs.CVcs.LGeess.IVarXiv:2001.01599v22020
  4. Instance-aware Image and Sentence Matching with Selective Multimodal LSTM

    Yan Huang, Wei Wang, Liang Wang

    cs.CVarXiv:1611.05588v12016
  5. ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

    Xionghao Wu, Yijun Yang, Shiyang Zhou +17

    cs.CVarXiv:2609.00188v12026
  6. LAFITE: Towards Language-Free Training for Text-to-Image Generation

    Yufan Zhou, Ruiyi Zhang, Changyou Chen +6

    cs.CVcs.LGarXiv:2111.13792v32021
  7. CSGNet: Neural Shape Parser for Constructive Solid Geometry

    Gopal Sharma, Rishabh Goyal, Difan Liu +2

    cs.CVcs.AIarXiv:1712.08290v22017
  8. SampleNet: Differentiable Point Cloud Sampling

    Itai Lang, Asaf Manor, Shai Avidan

    cs.CVarXiv:1912.03663v22019
  9. On the Reconstruction of Face Images from Deep Face Templates

    Guangcan Mai, Kai Cao, Pong C. Yuen +1

    cs.CVarXiv:1703.00832v42017
  10. H3-World: Turning Language Understanding into World Control

    Danze Chen, Zeqing Wang, Ziyue Lin +2

    cs.CVcs.AIarXiv:2609.01560v12026
  11. Deep Pixel-wise Binary Supervision for Face Presentation Attack Detection

    Anjith George, Sebastien Marcel

    cs.CVcs.CRarXiv:1907.04047v12019
  12. Cloze Test Helps: Effective Video Anomaly Detection via Learning to Complete Video Events

    Guang Yu, Siqi Wang, Zhiping Cai +4

    cs.CVcs.LGeess.IVarXiv:2008.11988v12020
  13. Document-level Relation Extraction as Semantic Segmentation

    Ningyu Zhang, Xiang Chen, Xin Xie +6

    cs.CLcs.AIcs.CVarXiv:2106.03618v22021
  14. Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

    Mingwang Xu, Hui Li, Qingkun Su +6

    cs.CVarXiv:2406.08801v22024
  15. H-vmunet: High-order Vision Mamba UNet for Medical Image Segmentation

    Renkai Wu, Yinghao Liu, Pengchen Liang +1

    cs.CVarXiv:2403.13642v12024
  16. Vision Language Models in Autonomous Driving: A Survey and Outlook

    Xingcheng Zhou, Mingyu Liu, Ekim Yurtsever +4

    cs.CVcs.AIarXiv:2310.14414v22023
  17. OPUS-V2: Bridging the Gap between Sparse Points and Dense Voxels

    Jiabao Wang, Qiang Meng, Liujiang Yan +3

    cs.CVarXiv:2608.29187v12026
  18. Convolutional Neural Fabrics

    Shreyas Saxena, Jakob Verbeek

    cs.CVcs.LGcs.NEarXiv:1606.02492v42016
  19. Plug and play methods for magnetic resonance imaging (long version)

    Rizwan Ahmad, Charles A. Bouman, Gregery T. Buzzard +4

    cs.CVarXiv:1903.08616v52019
  20. Towards Fully Automated Medical Imaging Code Generation via Validation-based Context Engineering

    Zixiao Zhao, Jing Sun, Zhe Hou +5

    cs.CVcs.SEarXiv:2608.29016v12026
  21. Audio-Visual Scene-Aware Dialog

    Huda Alamri, Vincent Cartillier, Abhishek Das +9

    cs.CVarXiv:1901.09107v22019
  22. RanPAC: Random Projections and Pre-trained Models for Continual Learning

    Mark D. McDonnell, Dong Gong, Amin Parveneh +2

    cs.LGcs.CVarXiv:2307.02251v32023
  23. PointDAN: A Multi-Scale 3D Domain Adaption Network for Point Cloud Representation

    Can Qin, Haoxuan You, Lichen Wang +2

    cs.CVcs.LGarXiv:1911.02744v12019
  24. EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

    Zhiyuan Chen, Jiajiong Cao, Zhiquan Chen +2

    cs.CVarXiv:2407.08136v22024
  25. Adaptive Multi-Teacher Multi-level Knowledge Distillation

    Yuang Liu, Wei Zhang, Jun Wang

    cs.CVarXiv:2103.04062v12021
  26. Cross-modality image synthesis from unpaired data using CycleGAN: Effects of gradient consistency loss and training data size

    Yuta Hiasa, Yoshito Otake, Masaki Takao +5

    cs.CVarXiv:1803.06629v32018
  27. ReFlowSET: Representation-Aligned Latent Flow Matching for SAR-to-EO Image Translation

    Jeonghyeok Do, Seungchul Lee, Munchurl Kim

    cs.CVarXiv:2609.00968v12026
  28. Cross-domain Contrastive Learning for Unsupervised Domain Adaptation

    Rui Wang, Zuxuan Wu, Zejia Weng +3

    cs.CVcs.AIcs.LGarXiv:2106.05528v22021
  29. CFC-Net: A Critical Feature Capturing Network for Arbitrary-Oriented Object Detection in Remote Sensing Images

    Qi Ming, Lingjuan Miao, Zhiqiang Zhou +1

    cs.CVarXiv:2101.06849v22021
  30. Automatic Radiology Report Generation based on Multi-view Image Fusion and Medical Concept Enrichment

    Jianbo Yuan, Haofu Liao, Rui Luo +1

    eess.IVcs.CVcs.MMarXiv:1907.09085v22019
  31. Learning monocular depth estimation infusing traditional stereo knowledge

    Fabio Tosi, Filippo Aleotti, Matteo Poggi +1

    cs.CVarXiv:1904.04144v12019
  32. Symbiotic Graph Neural Networks for 3D Skeleton-based Human Action Recognition and Motion Prediction

    Maosen Li, Siheng Chen, Xu Chen +3

    cs.CVarXiv:1910.02212v12019
  33. A General Framework for Adversarial Examples with Objectives

    Mahmood Sharif, Sruti Bhagavatula, Lujo Bauer +1

    cs.CVcs.CRarXiv:1801.00349v22017
  34. Memory-Attended Recurrent Network for Video Captioning

    Wenjie Pei, Jiyuan Zhang, Xiangrong Wang +3

    cs.CVarXiv:1905.03966v12019
  35. Segmentation of Bovid Dentition Under Imperfect Annotations: A Comparative Study of Convolutional and Attention Models

    Keith G. Mills, Evan B. Sanders, Gregory J. Matthews +1

    cs.CVcs.LGarXiv:2608.31052v12026
  36. Global Context Vision Transformers

    Ali Hatamizadeh, Hongxu Yin, Greg Heinrich +2

    cs.CVcs.AIcs.LGarXiv:2206.09959v52022
  37. Learn to Match: Automatic Matching Network Design for Visual Tracking

    Zhipeng Zhang, Yihao Liu, Xiao Wang +2

    cs.CVarXiv:2108.00803v12021
  38. Adaptive Sparse Convolutional Networks with Global Context Enhancement for Faster Object Detection on Drone Images

    Bowei Du, Yecheng Huang, Jiaxin Chen +1

    cs.CVarXiv:2303.14488v12023
  39. Marr Revisited: 2D-3D Alignment via Surface Normal Prediction

    Aayush Bansal, Bryan Russell, Abhinav Gupta

    cs.CVarXiv:1604.01347v12016
  40. Stable View Synthesis

    Gernot Riegler, Vladlen Koltun

    cs.CVarXiv:2011.07233v22020
  41. One Adapter, Many Tasks: Task-Conditioned Feature Transformations for Continual Learning

    Yunxiang Fu, Meng Lou, Yizhou Yu

    cs.CVcs.LGarXiv:2608.31096v12026
  42. On the Importance of Noise Scheduling for Diffusion Models

    Ting Chen

    cs.CVcs.GRcs.LGarXiv:2301.10972v42023
  43. Driving on Memory

    Christian Löwens, Thorben Funke, Alexandru Paul Condurache

    cs.CVcs.LGcs.ROarXiv:2608.31029v12026
  44. Zero-Shot Out-of-Distribution Detection Based on the Pre-trained Model CLIP

    Sepideh Esmaeilpour, Bing Liu, Eric Robertson +1

    cs.CVcs.LGarXiv:2109.02748v32021
  45. Findings of the Second Shared Task on Multimodal Machine Translation and Multilingual Image Description

    Desmond Elliott, Stella Frank, Loïc Barrault +2

    cs.CLcs.CVarXiv:1710.07177v12017
  46. UMDFaces: An Annotated Face Dataset for Training Deep Networks

    Ankan Bansal, Anirudh Nanduri, Carlos Castillo +2

    cs.CVarXiv:1611.01484v22016
  47. Orientation-boosted Voxel Nets for 3D Object Recognition

    Nima Sedaghat, Mohammadreza Zolfaghari, Ehsan Amiri +1

    cs.CVcs.NEarXiv:1604.03351v22016
  48. Deep Cross-Modal Audio-Visual Generation

    Lele Chen, Sudhanshu Srivastava, Zhiyao Duan +1

    cs.CVcs.MMcs.SDarXiv:1704.08292v12017
  49. Action Recognition Based on Joint Trajectory Maps with Convolutional Neural Networks

    Pichao Wang, Wanqing Li, Chuankun Li +1

    cs.CVarXiv:1612.09401v12016
  50. Learning JPEG Compression Artifacts for Image Manipulation Detection and Localization

    Myung-Joon Kwon, Seung-Hun Nam, In-Jae Yu +2

    eess.IVcs.CVcs.LGarXiv:2108.12947v22021
  51. Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks

    Sihyun Yu, Jihoon Tack, Sangwoo Mo +4

    cs.CVcs.LGarXiv:2202.10571v12022
  52. 3D Quasi-Recurrent Neural Network for Hyperspectral Image Denoising

    Kaixuan Wei, Ying Fu, Hua Huang

    cs.CVarXiv:2003.04547v12020
  53. Structured Sparse Subspace Clustering: A Joint Affinity Learning and Subspace Clustering Framework

    Chun-Guang Li, Chong You, René Vidal

    cs.CVarXiv:1610.05211v22016
  54. RAPIQUE: Rapid and Accurate Video Quality Prediction of User Generated Content

    Zhengzhong Tu, Xiangxu Yu, Yilin Wang +3

    cs.CVcs.MMeess.IVarXiv:2101.10955v22021
  55. Retina U-Net: Embarrassingly Simple Exploitation of Segmentation Supervision for Medical Object Detection

    Paul F. Jaeger, Simon A. A. Kohl, Sebastian Bickelhaupt +4

    cs.CVarXiv:1811.08661v12018
  56. STA: Spatial-Temporal Attention for Large-Scale Video-based Person Re-Identification

    Yang Fu, Xiaoyang Wang, Yunchao Wei +1

    cs.CVarXiv:1811.04129v12018
  57. MeteorNet: Deep Learning on Dynamic 3D Point Cloud Sequences

    Xingyu Liu, Mengyuan Yan, Jeannette Bohg

    cs.CVcs.LGcs.ROarXiv:1910.09165v22019
  58. [Extended version] Rethinking Deep Neural Network Ownership Verification: Embedding Passports to Defeat Ambiguity Attacks

    Lixin Fan, Kam Woh Ng, Chee Seng Chan

    cs.CRcs.CVcs.LGarXiv:1909.07830v32019
  59. Video Frame Interpolation Transformer

    Zhihao Shi, Xiangyu Xu, Xiaohong Liu +2

    cs.CVarXiv:2111.13817v32021
  60. Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data

    Chaoyi Wu, Xiaoman Zhang, Ya Zhang +2

    cs.CVcs.CLarXiv:2308.02463v52023