Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

9,121 to 9,180 of 18,817

  1. A Single Stream Network for Robust and Real-time RGB-D Salient Object Detection

    Xiaoqi Zhao, Lihe Zhang, Youwei Pang +2

    cs.CVarXiv:2007.06811v22020
  2. Who Said What: Modeling Individual Labelers Improves Classification

    Melody Y. Guan, Varun Gulshan, Andrew M. Dai +1

    cs.LGcs.CVarXiv:1703.08774v22017
  3. Transfer learning for music classification and regression tasks

    Keunwoo Choi, György Fazekas, Mark Sandler +1

    cs.CVcs.AIcs.MMarXiv:1703.09179v42017
  4. Revisiting Perspective Information for Efficient Crowd Counting

    Miaojing Shi, Zhaohui Yang, Chao Xu +1

    cs.CVarXiv:1807.01989v32018
  5. Contextual Diversity for Active Learning

    Sharat Agarwal, Himanshu Arora, Saket Anand +1

    cs.CVarXiv:2008.05723v12020
  6. Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-Identification

    Haowei Zhu, Wenjing Ke, Dong Li +3

    cs.CVcs.AIcs.LGarXiv:2205.02151v12022
  7. Vision Permutator: A Permutable MLP-Like Architecture for Visual Recognition

    Qibin Hou, Zihang Jiang, Li Yuan +3

    cs.CVarXiv:2106.12368v12021
  8. Resolution Adaptive Networks for Efficient Inference

    Le Yang, Yizeng Han, Xi Chen +3

    cs.CVarXiv:2003.07326v52020
  9. TCTrack: Temporal Contexts for Aerial Tracking

    Ziang Cao, Ziyuan Huang, Liang Pan +3

    cs.CVarXiv:2203.01885v32022
  10. Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions

    Jaewoo Ahn, Junseo Kim, Hyunseo Kim +4

    cs.CLcs.AIcs.CVarXiv:2608.30428v12026
  11. Learning to Generalize Unseen Domains via Memory-based Multi-Source Meta-Learning for Person Re-Identification

    Yuyang Zhao, Zhun Zhong, Fengxiang Yang +4

    cs.CVarXiv:2012.00417v32020
  12. SGUIE-Net: Semantic Attention Guided Underwater Image Enhancement with Multi-Scale Perception

    Qi Qi, Kunqian Li, Haiyong Zheng +3

    eess.IVcs.CVarXiv:2201.02832v12022
  13. CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions

    Tsung-Han Wu, Heekyung Lee, Anya Ji +4

    cs.CLcs.CVarXiv:2608.28958v12026
  14. Pose2Seg: Detection Free Human Instance Segmentation

    Song-Hai Zhang, Ruilong Li, Xin Dong +6

    cs.CVarXiv:1803.10683v32018
  15. Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching

    Rong Shan, Tianyi Xu, Congmin Zheng +9

    cs.CVcs.IRarXiv:2608.28695v12026
  16. Detecting Sarcasm in Multimodal Social Platforms

    Rossano Schifanella, Paloma de Juan, Joel Tetreault +1

    cs.CVcs.CLcs.MMarXiv:1608.02289v12016
  17. Time Will Tell: New Outlooks and A Baseline for Temporal Multi-View 3D Object Detection

    Jinhyung Park, Chenfeng Xu, Shijia Yang +4

    cs.CVcs.AIcs.LGarXiv:2210.02443v12022
  18. Domain Enhanced Arbitrary Image Style Transfer via Contrastive Learning

    Yuxin Zhang, Fan Tang, Weiming Dong +4

    cs.CVcs.GRarXiv:2205.09542v22022
  19. Understanding and Mitigating Copying in Diffusion Models

    Gowthami Somepalli, Vasu Singla, Micah Goldblum +2

    cs.LGcs.CRcs.CVarXiv:2305.20086v12023
  20. PiCIE: Unsupervised Semantic Segmentation using Invariance and Equivariance in Clustering

    Jang Hyun Cho, Utkarsh Mall, Kavita Bala +1

    cs.CVarXiv:2103.17070v12021
  21. Progressive Transformers for End-to-End Sign Language Production

    Ben Saunders, Necati Cihan Camgoz, Richard Bowden

    cs.CVcs.CLcs.LGarXiv:2004.14874v22020
  22. Will we run out of data? Limits of LLM scaling based on human-generated data

    Pablo Villalobos, Anson Ho, Jaime Sevilla +3

    cs.LGcs.AIcs.CLarXiv:2211.04325v22022
  23. Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal Fusion

    Sijie Mai, Haifeng Hu, Songlong Xing

    cs.CVcs.LGcs.MMarXiv:1911.07848v42019
  24. Wave-ViT: Unifying Wavelet and Transformers for Visual Representation Learning

    Ting Yao, Yingwei Pan, Yehao Li +2

    cs.CVcs.LGarXiv:2207.04978v12022
  25. Learning Human Motion Models for Long-term Predictions

    Partha Ghosh, Jie Song, Emre Aksan +1

    cs.CVarXiv:1704.02827v22017
  26. Text-To-4D Dynamic Scene Generation

    Uriel Singer, Shelly Sheynin, Adam Polyak +8

    cs.CVcs.AIcs.LGarXiv:2301.11280v12023
  27. Knowledge Adaptation for Efficient Semantic Segmentation

    Tong He, Chunhua Shen, Zhi Tian +3

    cs.CVarXiv:1903.04688v12019
  28. DeepTAM: Deep Tracking and Mapping

    Huizhong Zhou, Benjamin Ummenhofer, Thomas Brox

    cs.CVarXiv:1808.01900v22018
  29. The Feeling of Success: Does Touch Sensing Help Predict Grasp Outcomes?

    Roberto Calandra, Andrew Owens, Manu Upadhyaya +4

    cs.ROcs.CVcs.LGarXiv:1710.05512v22017
  30. Global canopy height regression and uncertainty estimation from GEDI LIDAR waveforms with deep ensembles

    Nico Lang, Nikolai Kalischek, John Armston +3

    cs.LGcs.CVarXiv:2103.03975v22021
  31. Generative replay with feedback connections as a general strategy for continual learning

    Gido M. van de Ven, Andreas S. Tolias

    cs.LGcs.AIcs.CVarXiv:1809.10635v22018
  32. Towards Effective Low-bitwidth Convolutional Neural Networks

    Bohan Zhuang, Chunhua Shen, Mingkui Tan +2

    cs.CVarXiv:1711.00205v22017
  33. ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

    Anoop Senthil

    cs.CLcs.CVarXiv:2608.28707v12026
  34. RMT: Retentive Networks Meet Vision Transformers

    Qihang Fan, Huaibo Huang, Mingrui Chen +2

    cs.CVarXiv:2309.11523v62023
  35. On learning to localize objects with minimal supervision

    Hyun Oh Song, Ross Girshick, Stefanie Jegelka +3

    cs.CVcs.LGarXiv:1403.1024v42014
  36. Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners

    Renrui Zhang, Xiangfei Hu, Bohao Li +5

    cs.CVcs.CLarXiv:2303.02151v12023
  37. Person Re-Identification via Recurrent Feature Aggregation

    Yichao Yan, Bingbing Ni, Zhichao Song +3

    cs.CVarXiv:1701.06351v12017
  38. Learning a Text-Video Embedding from Incomplete and Heterogeneous Data

    Antoine Miech, Ivan Laptev, Josef Sivic

    cs.CVarXiv:1804.02516v22018
  39. SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian Splatting

    Zhijing Shao, Zhaolong Wang, Zhuang Li +5

    cs.GRcs.CVarXiv:2403.05087v12024
  40. Embedding Structured Contour and Location Prior in Siamesed Fully Convolutional Networks for Road Detection

    Qi Wang, Junyu Gao, Yuan Yuan

    cs.CVarXiv:1905.01575v12019
  41. DC-UNet: Rethinking the U-Net Architecture with Dual Channel Efficient CNN for Medical Images Segmentation

    Ange Lou, Shuyue Guan, Murray Loew

    eess.IVcs.CVarXiv:2006.00414v12020
  42. Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory

    Runjia Qian, Zile Wang, Jihai Zhang +14

    cs.CVarXiv:2608.29910v12026
  43. RGBT Salient Object Detection: A Large-scale Dataset and Benchmark

    Zhengzheng Tu, Yan Ma, Zhun Li +3

    cs.CVarXiv:2007.03262v62020
  44. Object Motion Guided Human Motion Synthesis

    Jiaman Li, Jiajun Wu, C. Karen Liu

    cs.CVarXiv:2309.16237v12023
  45. OmniDepth: Dense Depth Estimation for Indoors Spherical Panoramas

    Nikolaos Zioulis, Antonis Karakottas, Dimitrios Zarpalas +1

    cs.CVarXiv:1807.09620v12018
  46. Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection

    Huan Liu, Zichang Tan, Chuangchuang Tan +3

    cs.CVarXiv:2312.16649v12023
  47. Abdominal multi-organ segmentation with organ-attention networks and statistical fusion

    Yan Wang, Yuyin Zhou, Wei Shen +3

    cs.CVarXiv:1804.08414v12018
  48. Explainable Medical Imaging AI Needs Human-Centered Design: Guidelines and Evidence from a Systematic Review

    Haomin Chen, Catalina Gomez, Chien-Ming Huang +1

    cs.HCcs.CVcs.LGarXiv:2112.12596v42021
  49. Practical Blind Image Denoising via Swin-Conv-UNet and Data Synthesis

    Kai Zhang, Yawei Li, Jingyun Liang +6

    cs.CVcs.GReess.IVarXiv:2203.13278v42022
  50. Parametric Multimodal User Memory: Storing What Captions Cannot Carry

    Bojie Li, Noah Shi

    cs.CLcs.AIcs.CVarXiv:2608.28609v12026
  51. Batch-Instance Normalization for Adaptively Style-Invariant Neural Networks

    Hyeonseob Nam, Hyo-Eun Kim

    cs.CVarXiv:1805.07925v32018
  52. DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

    Jiashu Zhu, Yanhao Zheng, Ruitian Tian +7

    cs.CVcs.SDarXiv:2608.31106v12026
  53. Fast Inference in Sparse Coding Algorithms with Applications to Object Recognition

    Koray Kavukcuoglu, Marc'Aurelio Ranzato, Yann LeCun

    cs.CVcs.LGarXiv:1010.3467v12010
  54. Defense Against Adversarial Attacks Using Feature Scattering-based Adversarial Training

    Haichao Zhang, Jianyu Wang

    cs.CVcs.CRcs.LGarXiv:1907.10764v42019
  55. Learning Task-Oriented Grasping for Tool Manipulation from Simulated Self-Supervision

    Kuan Fang, Yuke Zhu, Animesh Garg +4

    cs.ROcs.CVcs.LGarXiv:1806.09266v12018
  56. Image Deformation Meta-Networks for One-Shot Learning

    Zitian Chen, Yanwei Fu, Yu-Xiong Wang +3

    cs.CVarXiv:1905.11641v22019
  57. Shield: Fast, Practical Defense and Vaccination for Deep Learning using JPEG Compression

    Nilaksh Das, Madhuri Shanbhogue, Shang-Tse Chen +5

    cs.CVcs.AIcs.CRarXiv:1802.06816v12018
  58. Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained Models

    Guillermo Ortiz-Jimenez, Alessandro Favero, Pascal Frossard

    cs.LGcs.CVarXiv:2305.12827v32023
  59. Unsupervised Domain Adaptation through Self-Supervision

    Yu Sun, Eric Tzeng, Trevor Darrell +1

    cs.LGcs.CVstat.MLarXiv:1909.11825v22019
  60. HyperSeg: Patch-wise Hypernetwork for Real-time Semantic Segmentation

    Yuval Nirkin, Lior Wolf, Tal Hassner

    cs.CVarXiv:2012.11582v22020