Keyframe-Based Multimodal Video Archive Retrieval Using CLIP Embeddings: A Scalable and Secure Framework for NHK Archives
Large-scale video archives, such as the NHK Archives with over ten million assets, face critical retrieval challenges: manual metadata is costly and incomplete. At the same time, cloud-based AI solutions introduce unacceptable data security risks. This paper presents AXIS (Archives Cross-modal Intelligent Search), a metadata-free retrieval system built on three integrated components. First, keyframe-centric processing reduces the analysis corpus by 100-1,000× versus full-frame processing–making billion-scale analysis computationally tractable. Second, offline deployment of a Japanese-language Contrastive Language-Image Pretraining (CLIP) model enables on-premise semantic feature extraction, ensuring data sovereignty. Third, Approximate Nearest Neighbor (ANN) indexing via Facebook AI Smiliary Search (FAISS) enables sub-second retrieval at scale. Evaluated on a 30-million-keyframe subset of the NHK Archives, the system achieves a mean query latency of 0.48 sec (text) and 0.64 sec (image), processing 8.4 million keyframes per day on commodity CPU hardware. The hybrid on-premise/cloud architecture supports intuitive multimodal search while preserving institutional data security.
- Print ISSN
- 1545-0279
- Electronic ISSN
- 2160-2492
- Published
- 2026-09
- Content type
- Original Research
- Keywords
- nhk archives, content retrieval, keyframe-centric processing, clip (contrastive language-image pre-training)
- DOI
- 10.5594/JMI.2026/CJWA9953