Skip to yearly menu bar Skip to main content

Invited Talk
Affinity Workshop: Muslims in ML

Data Paucity and Low Resource Scenarios: Challenges and Opportunities

Mona Diab


In an era unstructured data abundance, you would think that we have solved our data requirements for building robust systems for language processing. However, this is not the case if we think on a global scale with over 7000 languages where only a handful have digital resources. Moreover, systems at scale with good performance typically require annotated resources.The existence of a handful of resources in a some languages is a reflection of the digital disparity in various societies leading to inadvertent biases in systems. In this talk I will show some solutions for low resource scenarios, both cross domain and genres as well as cross lingually.