CAFE: Spontaneous code-switching speech dataset in Algerian dialect, French and English

CAFE：阿尔及利亚方言、法语和英语的自发语码转换语音数据集

阅读：2

作者：Lachemat,Houssam Eddine-Othman,Akli,Abbas,Oukas,Nourredine,El Kheir,Yassine,Haboussi,Samia,Chowdhury,Shammur Absar

期刊：	Data in Brief	影响因子：	1.400
时间：	2025	起止号：	2025 Dec;63:112150
doi：	10.1016/j.dib.2025.112150

Abstract

Publicly available datasets capturing spontaneous multilingual speech-especially those involving code-switching between Algerian Arabic, French, and English-are critically scarce. This lack of resources hinders the development of automatic speech recognition (ASR) and multilingual NLP systems for low-resource languages and under-represented Arabic dialects. We introduce CAFE, a novel dataset comprising approximately 37 h of spontaneous, in vivo human-human conversations among 100+ speakers across Algeria. The dialogues cover diverse everyday topics such as sports, science, and technology, and exhibit a rich range of natural conversational phenomena, including explicit code-switching, overlapping speech, non-lexical vocalizations (e.g., laughter, fillers, ambient noise), and dialectal variation reflecting Algeria's sociolinguistic landscape. CAFE is released in two tiers: CAFE-small (2 h 36 m): A fully human-annotated subset with high-quality transcriptions, vocal event labels, and linguistic annotations, supporting ASR evaluation, NLP tasks, and code-switching analysis. CAFE-large (∼34 h 35 m): The remainder of the corpus, automatically labeled, suitable for pretraining and semi-supervised learning. To support controlled experiments, CAFE-small includes two curated subsets: (i) CAFE-small-clean (2 h 18 m): Contains utterances with no overlapping speech. (ii) CAFE-small-overlap (17 m): Contains 23 files with overlap segments and timestamps. The dataset also provides rich metadata, including audio chunk IDs, dialect labels, and both raw and linguistically processed transcripts. CAFE offers a valuable resource for advancing ASR, dialect identification, and sociolinguistic analysis in multilingual and low-resource settings.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用；引用内容仅为补充信息，不代表本站立场。

2、若认为本页面引用内容涉及侵权，请及时与本站联系，我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容，需注明“来源：[生知库]”并获得授权；使用引用内容的，需自行联系原作者获得许可。

4、投稿及合作请联系：info@biocloudy.com。

肿瘤免疫

炎症

T细胞

线粒体

凋亡

转录调控

巨噬细胞

自噬

传染病

氧化应激

肠道菌群

磷酸化

囊泡

血管生成

3D/类器官

单细胞

中性粒细胞

外泌体

DNA甲基化

miRNA

药物研究

铁死亡

细胞衰老

乙酰化

缺氧低氧

泛素化

树突状细胞

炎性小体

肿瘤微环境

组蛋白修饰

lncRNA

代谢重编程

焦亡

m6A/m5C/m7G

内质网应激

空间多组学

细胞基因治疗

相分离

治疗耐药

Treg

上皮间质转化

免疫代谢

染色质重塑

脂质过氧化

蛋白质稳态

脂代谢

铁代谢

细胞极性

氨基酸代谢

cGAS-STING

碱基编辑

蛋白降解

肠脑轴

翻译调控

乳酸化

circRNA

piRNA

肿瘤异质性

NK 细胞

氧化脂质

MDSC

NETosis

溶酶体功能

低氧缺氧

琥珀酰化

细胞干性

CAR-NK

冷应激

RNA 编辑

Tfh

巴豆酰化

器官芯片

器官纤维化

表观遗传记忆

铜死亡

线粒体未折叠蛋白反应

空间代谢组

程序性坏死

自噬流

丙酰化

MAIT 细胞

肠肝轴