Keyframe-Based Multimodal Video Archive Retrieval Using CLIP Embeddings: A Scalable and Secure Framework for NHK Archives

Yo Narita

Large-scale video archives, such as the NHK Archives with over ten million assets, face critical retrieval challenges: manual metadata is costly and incomplete. At the same time, cloud-based AI solutions introduce unacceptable data security risks. This paper presents AXIS (Archives Cross-modal Intelligent Search), a metadata-free retrieval system built on three integrated components. First, keyframe-centric processing reduces the analysis corpus by 100-1,000× versus full-frame processing–making billion-scale analysis computationally tractable. Second, offline deployment of a Japanese-language Contrastive Language-Image Pretraining (CLIP) model enables on-premise semantic feature extraction, ensuring data sovereignty. Third, Approximate Nearest Neighbor (ANN) indexing via Facebook AI Smiliary Search (FAISS) enables sub-second retrieval at scale. Evaluated on a 30-million-keyframe subset of the NHK Archives, the system achieves a mean query latency of 0.48 sec (text) and 0.64 sec (image), processing 8.4 million keyframes per day on commodity CPU hardware. The hybrid on-premise/cloud architecture supports intuitive multimodal search while preserving institutional data security.

Print ISSN
Electronic ISSN
2160-2492
Published
2026-09
Content type
Original Research
Keywords
nhk archives, content retrieval, keyframe-centric processing, clip (contrastive language-image pre-training)
DOI
10.5594/JMI.2026/CJWA9953