Introduction
The OU-CA: A Face-Voice Multimodal Benchmark for Continuous User Authentication is meant to aid research efforts in the general area of developing, testing and evaluating algorithms for continuous authentication using face and voice. The D3 Center, the University of Osaka (OU) has copyright in the collection of face video, voice and associated data and serves as a distributor of the OU-CA dataset.
The database was collected through a web-based conversational game conducted between pairs of participants. To minimize the influence of background noise or environmental interference, we prepared two sound-isolated booths, where each participant sat individually and interacted with their partner via desktop computers connected through an online meeting. To capture facial videos with diverse pose variations, we employed a custom-designed display system equipped with five synchronized cameras positioned at the top-left, bottomleft, bottom-center, top-right, and bottom-right corners, as shown in Fig. 1. All cameras were of the same model (i.e., TRI023s-CC camera) and synchronized using a hardware synchronization signal generator. The captured facial images were saved in “BayerRG” format with a resolution of 1920×1200 pixels and a frame rate of 7.5 fps.
Fig. 1: (a) Capturing setup with five example faces and (b) The used display with five face cameras marked by red circles.
The conversational interaction was designed in the form of two Japanese-style word games: “Shiritori” (word-chain) and “Rensou” (word association). In Shiritori, each participant says a word beginning with the final syllable of their partner’s previous word, while in Rensou, participants respond with semantically related words. These game formats were selected to naturally evoke spontaneous voice and facial expressions during interaction. Each game session lasted approximately five minutes, allowing participants to engage in free conversation without any content restrictions. The audio data was recorded using a condenser microphone (i.e., Audio-Technica AT2020) and stored in WAV format with a bit rate of 1,411 kbps.
The detailed descriptions are found in the following paper.
- X. Li, C. Xu, Y. Yagi, "OU-CA: A Face-Voice Multimodal Benchmark for Continuous User Authentication", IEEE Transactions on Biometrics, Behavior, and Identity Science (T-BIOM), 2026. [Bib]
Dataset
The collected face and voice data are from 1,169 volunteers (331 males and 838 females) with ages ranging from 7 to 82 years old. The detailed age and gender distributions are illustrated in Fig. 2(a). Each participant took part in two conversational game sessions, during which facial videos were captured simultaneously by five cameras, resulting in up to ten face sequences per subject. Due to occasional missed captures from a small number of cameras, a few subjects have fewer than ten sequences. In total, the database contains 11,650 face sequences. Note that some sequences may contain fewer frames (fewer than 2,200 frames) due to early termination of the capture. The distribution of frame counts per sequence is shown in Fig. 2(b).
Fig. 2: Statistics of the proposed OU-CA. (a) Distribution of subjects’ age and gender. (b) Distribution of the number of frames in each sequence.
For voice, however, the recorded conversational data often consists of short phrases or isolated words, making it less suitable for enrollment. Therefore, we additionally collected voice enrollment data by asking each subject to read a fixed Japanese script. The recorded audio lasts approximately 60 seconds and is saved in the same format as the probe data.
How to get the dataset?
For ease of use, this dataset provides cropped face images. For details of the preprocessing procedure, please refer to the original paper. This dataset, including a set of size-normalized cropped faces (112x112) and voice recordings , could be downloaded as a zip file with password protection and the password will be issued on a case-by-case basis. To receive the password, the requestor must send the release agreement signed by a legal representative of your institution (e.g., your supervisor if you are a student) to the database administrator by mail, e-mail, or FAX.- Release agreement
- Dataset (The download URL and password will be provided once your application is approved.)
The database administrator
Yagi Lab., D3 Center, The University of OsakaAddress: 8-1 Mihogaoka, Ibaraki, Osaka, 567-0047, JAPAN
FAX: +81-6-6879-4032.