Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

8,461 to 8,520 of 18,866

  1. TCLR: Temporal Contrastive Learning for Video Representation

    Ishan Dave, Rohit Gupta, Mamshad Nayeem Rizve +1

    cs.CVarXiv:2101.07974v42021
  2. A Unified Feature Disentangler for Multi-Domain Image Translation and Manipulation

    Alexander H. Liu, Yen-Cheng Liu, Yu-Ying Yeh +1

    cs.CVarXiv:1809.01361v32018
  3. Intra-Inter Camera Similarity for Unsupervised Person Re-Identification

    Shiyu Xuan, Shiliang Zhang

    cs.CVcs.AIarXiv:2103.11658v12021
  4. Beyond Visual Boundaries: Rethinking Scene Segmentation for Movie RAG

    Dong-Hee Kim, Seonwoo Choi, Changbeen Kim +6

    cs.CVarXiv:2608.28699v12026
  5. Better Aggregation in Test-Time Augmentation

    Divya Shanmugam, Davis Blalock, Guha Balakrishnan +1

    cs.CVarXiv:2011.11156v22020
  6. SCoPE-Reg: Efficient Rigid Ultrasound Slice-to-Volume Registration via State-Space Correlation and Closed-Form Pose Estimation

    Niklas Schwarz, Jens Kleesiek, Moritz Rempe

    eess.IVcs.CVarXiv:2608.28715v12026
  7. Visually Grounded Reasoning across Languages and Cultures

    Fangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti +3

    cs.CLcs.AIcs.CVarXiv:2109.13238v22021
  8. mmIR: Frequency-Space Inverse Rendering for 3D Millimeter-Wave Radar ADC Synthesis

    Adnan Armouti, Yixuan Gao, Rajalakshmi Nandakumar

    cs.CVcs.GRcs.LGarXiv:2608.28913v12026
  9. ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement

    Jinlong Li, Jiaming Ding, Dingfu Lu +8

    cs.CVarXiv:2608.28784v12026
  10. Seamless Scene Segmentation

    Lorenzo Porzi, Samuel Rota Bulò, Aleksander Colovic +1

    cs.CVarXiv:1905.01220v12019
  11. Towards Multi-pose Guided Virtual Try-on Network

    Haoye Dong, Xiaodan Liang, Bochao Wang +3

    cs.CVarXiv:1902.11026v12019
  12. Reliability Does Matter: An End-to-End Weakly Supervised Semantic Segmentation Approach

    Bingfeng Zhang, Jimin Xiao, Yunchao Wei +2

    cs.CVarXiv:1911.08039v12019
  13. Stereo Correspondence and Reconstruction of Endoscopic Data Challenge

    Max Allan, Jonathan Mcleod, Congcong Wang +21

    cs.CVarXiv:2101.01133v42021
  14. Augmented Skeleton Based Contrastive Action Learning with Momentum LSTM for Unsupervised Action Recognition

    Haocong Rao, Shihao Xu, Xiping Hu +2

    cs.CVarXiv:2008.00188v42020
  15. TableBank: A Benchmark Dataset for Table Detection and Recognition

    Minghao Li, Lei Cui, Shaohan Huang +3

    cs.CVarXiv:1903.01949v22019
  16. PYSKL: Towards Good Practices for Skeleton Action Recognition

    Haodong Duan, Jiaqi Wang, Kai Chen +1

    cs.CVarXiv:2205.09443v12022
  17. Visual Alignment Constraint for Continuous Sign Language Recognition

    Yuecong Min, Aiming Hao, Xiujuan Chai +1

    cs.CVcs.HCarXiv:2104.02330v22021
  18. SemanticPOSS: A Point Cloud Dataset with Large Quantity of Dynamic Instances

    Yancheng Pan, Biao Gao, Jilin Mei +3

    cs.ROcs.CVeess.IVarXiv:2002.09147v12020
  19. Detecting the Unexpected via Image Resynthesis

    Krzysztof Lis, Krishna Nakka, Pascal Fua +1

    cs.CVarXiv:1904.07595v22019
  20. Survey on Deep Neural Networks in Speech and Vision Systems

    Mahbubul Alam, Manar D. Samad, Lasitha Vidyaratne +2

    cs.CVcs.LGcs.NEarXiv:1908.07656v22019
  21. Contrastive Attention for Automatic Chest X-ray Report Generation

    Fenglin Liu, Changchang Yin, Xian Wu +5

    cs.CVcs.CLarXiv:2106.06965v52021
  22. Efficient Self-supervised Vision Transformers for Representation Learning

    Chunyuan Li, Jianwei Yang, Pengchuan Zhang +5

    cs.CVcs.AIcs.LGarXiv:2106.09785v22021
  23. A Robust Volumetric Transformer for Accurate 3D Tumor Segmentation

    Himashi Peiris, Munawar Hayat, Zhaolin Chen +2

    eess.IVcs.CVarXiv:2111.13300v22021
  24. Collaborative Learning for Deep Neural Networks

    Guocong Song, Wei Chai

    stat.MLcs.CVcs.LGarXiv:1805.11761v22018
  25. Dataset Distillation using Neural Feature Regression

    Yongchao Zhou, Ehsan Nezhadarya, Jimmy Ba

    cs.LGcs.CVarXiv:2206.00719v22022
  26. SNF-Bench: Separating Static Drift from Natural Flow in Long-Horizon Fixed-Camera Video Generation

    Matiur Rahman Minar, Seunghun Oh, Ganghyeon Jeong +1

    cs.CVarXiv:2608.28694v12026
  27. Use of a Capsule Network to Detect Fake Images and Videos

    Huy H. Nguyen, Junichi Yamagishi, Isao Echizen

    cs.CVarXiv:1910.12467v22019
  28. Self-Guided and Cross-Guided Learning for Few-Shot Segmentation

    Bingfeng Zhang, Jimin Xiao, Terry Qin

    cs.CVarXiv:2103.16129v12021
  29. Restricting the Flow: Information Bottlenecks for Attribution

    Karl Schulz, Leon Sixt, Federico Tombari +1

    stat.MLcs.CVcs.LGarXiv:2001.00396v42020
  30. XFeat: Accelerated Features for Lightweight Image Matching

    Guilherme Potje, Felipe Cadar, Andre Araujo +2

    cs.CVarXiv:2404.19174v12024
  31. SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera Videos

    Haisong Liu, Yao Teng, Tao Lu +2

    cs.CVarXiv:2308.09244v22023
  32. Towards Multi-spatiotemporal-scale Generalized PDE Modeling

    Jayesh K. Gupta, Johannes Brandstetter

    cs.LGcs.CVarXiv:2209.15616v22022
  33. Constrained Deep Weak Supervision for Histopathology Image Segmentation

    Zhipeng Jia, Xingyi Huang, Eric I-Chao Chang +1

    cs.CVarXiv:1701.00794v12017
  34. A Generative Model of People in Clothing

    Christoph Lassner, Gerard Pons-Moll, Peter V. Gehler

    cs.CVarXiv:1705.04098v32017
  35. Coronary Mask Guided Registration for Continuous Time 4D Cardiac CT Dataset Construction

    Yuang Wang, Shuo Wang, Changyu Chen +12

    eess.IVcs.CVarXiv:2608.28712v12026
  36. Towards Accurate Markerless Human Shape and Pose Estimation over Time

    Yinghao Huang, Federica Bogo, Christoph Lassner +4

    cs.CVarXiv:1707.07548v52017
  37. Compact Snapshot Spectral Imaging with Calibration-Free Aperture Diffraction

    Tao Lv, Quan Yuan, Shiqiao Li +5

    cs.CVarXiv:2608.29230v12026
  38. Pixel-wise Geo-registration of Drone and Satellite Images

    Qingyang Liu, David G Shatwell, Parth Parag Kulkarni +1

    cs.CVarXiv:2608.28891v12026
  39. From CNN to Transformer: A Review of Medical Image Segmentation Models

    Wenjian Yao, Jiajun Bai, Wei Liao +3

    eess.IVcs.CVcs.LGarXiv:2308.05305v12023
  40. Do Convolutional Neural Networks Learn Class Hierarchy?

    Bilal Alsallakh, Amin Jourabloo, Mao Ye +2

    cs.CVarXiv:1710.06501v12017
  41. Elastic Triangle Splatting

    Tian Shi, Shenhan Qian, Daniel Cremers

    cs.CVcs.CGarXiv:2608.29106v12026
  42. 3D Ken Burns Effect from a Single Image

    Simon Niklaus, Long Mai, Jimei Yang +1

    cs.CVcs.GRarXiv:1909.05483v12019
  43. Uncertainty-aware Short-term Motion Prediction of Traffic Actors for Autonomous Driving

    Nemanja Djuric, Vladan Radosavljevic, Henggang Cui +5

    cs.LGcs.CVcs.ROarXiv:1808.05819v32018
  44. Adversarial Calibration Attack on Autonomous Vehicles

    Liangkai Liu, Qingzhao Zhang, Kang G. Shin

    cs.ROcs.CVcs.ETarXiv:2608.28778v12026
  45. Pothole Detection Based on Disparity Transformation and Road Surface Modeling

    Rui Fan, Umar Ozgunalp, Brett Hosking +2

    cs.CVeess.IVarXiv:1908.00894v32019
  46. EMMA: End-to-End Multimodal Model for Autonomous Driving

    Jyh-Jing Hwang, Runsheng Xu, Hubert Lin +11

    cs.CVcs.AIcs.CLarXiv:2410.23262v32024
  47. Real-Time Intensity-Image Reconstruction for Event Cameras Using Manifold Regularisation

    Christian Reinbacher, Gottfried Graber, Thomas Pock

    cs.CVarXiv:1607.06283v22016
  48. NerfDiff: Single-image View Synthesis with NeRF-guided Distillation from 3D-aware Diffusion

    Jiatao Gu, Alex Trevithick, Kai-En Lin +4

    cs.CVcs.LGarXiv:2302.10109v12023
  49. Discriminative Adversarial Domain Adaptation

    Hui Tang, Kui Jia

    cs.CVcs.LGarXiv:1911.12036v22019
  50. PersonNet: Person Re-identification with Deep Convolutional Neural Networks

    Lin Wu, Chunhua Shen, Anton van den Hengel

    cs.CVarXiv:1601.07255v22016
  51. FAMNet: Joint Learning of Feature, Affinity and Multi-dimensional Assignment for Online Multiple Object Tracking

    Peng Chu, Haibin Ling

    cs.CVarXiv:1904.04989v12019
  52. UI-Venus-2 Technical Report

    Venus Team, Zhuohan Cai, Haoxing Chen +28

    cs.AIcs.CLcs.CVarXiv:2609.00028v12026
  53. Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

    Xin Zhou, Zongchuang Zhao, Zhibo Yang +13

    cs.CVarXiv:2609.00111v12026
  54. CIAN: Cross-Image Affinity Net for Weakly Supervised Semantic Segmentation

    Junsong Fan, Zhaoxiang Zhang, Tieniu Tan +2

    cs.CVarXiv:1811.10842v22018
  55. T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval

    Xiaohan Wang, Linchao Zhu, Yi Yang

    cs.CVcs.MMarXiv:2104.10054v12021
  56. Likelihood Regret: An Out-of-Distribution Detection Score For Variational Auto-encoder

    Zhisheng Xiao, Qing Yan, Yali Amit

    cs.LGcs.CVstat.MLarXiv:2003.02977v32020
  57. Unsupervised Sparse Dirichlet-Net for Hyperspectral Image Super-Resolution

    Ying Qu, Hairong Qi, Chiman Kwan

    cs.CVarXiv:1804.05042v32018
  58. MULLS: Versatile LiDAR SLAM via Multi-metric Linear Least Square

    Yue Pan, Pengchuan Xiao, Yujie He +2

    cs.ROcs.CVarXiv:2102.03771v32021
  59. Decoupling Features in Hierarchical Propagation for Video Object Segmentation

    Zongxin Yang, Yi Yang

    cs.CVarXiv:2210.09782v32022
  60. Learning from Extrinsic and Intrinsic Supervisions for Domain Generalization

    Shujun Wang, Lequan Yu, Caizi Li +2

    cs.CVarXiv:2007.09316v12020