Mexican Spanish viseme phoneme /be/ dataset
Description
This dataset contains the dynamic performance of the Mexican Spanish phoneme /be/ by 10 people repeated three times by everyone. Multimodal data was captured using a MS Kinect V2 sensor connected to a laptop. RGB images (24 bits per pixel) and depth images (16 bit per pixel) sequences were captured at 30 frames per second. The synchronized audio of the phoneme was recorded in a WAV file at 44 Khz, 16 bits per sample. For each phoneme performance the following data are available for the mouth region: 90, 256x256 RGB images (24 bits per pixel) in PNG format 90, 256x256 depth images (16 bits per pixel) in PNG format One WAV file at 44 Khz, 16 bits per pixel To facilitate analysis, processing, machine learning and inference methods to be developed the following additional files are provided which are derived from the above raw data: 191 facial landmarks (x,y) extracted with MediaPipe software for each RGB image. The landmarks are stored in a TXT file, with one point (x,y) in each row. 191 facial landmarks (x,y,Z) extracted with MediaPipe software for each depth image. The landmarks are stored in a TXT file, with one point (x,y,Z) in each row. Cloud point (X,Y,Z) corresponding to each depth image obtained from reprojection using the depth camera intrinsic parameters. The cloud is stored as a TXT file, with one point (X,Y,Z) per row. Cloud point (X,Y,Z) corresponding to 191 landmarks of each depth image obtained from reprojection using the depth camera intrinsic parameters. The cloud is stored as a TXT file, with one point (X,Y,Z) per row.
Files
Steps to reproduce
The study focused on 21 phonemes of Mexican Spanish: A, BE, CHE, DA, E, FA, GA, I, JA, KA, LA, MA, ÑA, NO, O, PA, RA, SI, TE, U, and YA. Each of these phonemes data was organized into folders named FONEMA_*, where the asterisk is replaced by the respective phoneme. In each of these folders, you will find subfolders corresponding to the 10 participants (P1, P2, …, P10). Within each participant’s folder are subfolders corresponding to the three repetitions that were recorded (labeled as ENSAYO1, ENSAYO2, and ENSAYO3). And within each of these “ENSAYO*” (ESSAY) subfolders, the following information is available: captura_3s_audio.wav: This is the audio file captured by the Kinect of the phoneme being pronounced. Each one weighs approximately 517 KB. landmarks: This is a folder containing two subfolders with landmarks extracted from RGB images (each file is 2.84 KB in size, and the entire folder is 245 KB) and depth images (each file is 4.24 KB in size, and the entire folder is 367 KB) of the image cropped to 256 x 256 pixels. The “x” and “y” coordinates range from 0 to 255, and the Z value in the depth image is in mm. The depth images use depth landmarks extracted from the image that has already been processed using the median filter after outliers were corrected. nube_puntos: This is a folder containing 90 .txt files with the “X,” “Y,” and “Z” coordinates of each of the captured frames, allowing them to be viewed in three dimensions; each file is 2.21 MB in size. It also contains a subfolder with the point clouds of the reference points obtained through reprojection. These landmarks were obtained by cropping the region of interest from the original image; each file is 4.74 KB in size. This folder takes up 200 MB of storage space. overlay: This is a folder containing 90 images in RGB and depth PNG formats (side by side) with landmarks overlaid to show what they look like and which landmarks are extracted from the cropped images using MediaPipe. Each of these images is __ in size, and each folder is __ in size. recortes_boca: This is a folder containing 180 images in PNG format of the mouth region: 90 RGB images and 90 depth images. These images have been cropped to 256 x 256 pixels, because the RGB images were captured at 1920 x 1080 pixels and the depth images at 512 x 424 pixels; therefore, we cropped them to a size that can be processed by most neural networks architectures—in this case, 256 x 256 pixels. Depth images have been processed using a median filter to remove any outliers that may exist in the image. Each RGB image is 52 KB, each depth image is 14 KB, and each of these folders totals approximately 5.8 MB.
Institutions
- Universidad VeracruzanaVeracruz, Xalapa