Sign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities, yet advances in sign-language recognition, translation, and production remain constrained by fragmented datasets, inconsistent annotations, and limited linguistic coverage. A new survey, published on arXiv, provides a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages, and systematically analyzes key challenges such as modality imbalance, annotation granularity, and signer bias.
According to the survey by researchers Ni, Yiming; Cheng, Zhi-Qi; Li, Jiayu; and Wei, existing benchmarks often fail to reflect real-world communication needs. The work outlines considerations for future dataset design and introduces a 24-field Sign-Language Datasheet to support standardized documentation. The authors also released a public GitHub repository to enable reproducible evaluation.
Survey Scope and Scale
The study catalogs 120 distinct sign-language datasets spanning 35 different sign languages. This breadth highlights the diversity of resources available but also underscores fragmentation. The researchers note that while some languages like American Sign Language (ASL) and Chinese Sign Language (CSL) are well-represented, many others have few or no large-scale datasets.
| Metric | Value |
|---|---|
| Number of datasets | 120 |
| Sign languages covered | 35 |
| Datasheet fields | 24 |
| Repository | Public GitHub |
Key Challenges Identified
The survey identifies three primary challenges:
- Modality Imbalance: Most datasets rely on single camera views or limited sensor modalities, failing to capture the full expressiveness of signing.
- Annotation Granularity: Annotations vary widely — from coarse gloss labels to precise linguistic features — making cross-dataset comparisons difficult.
- Signer Bias: Many datasets include a small number of signers, leading to models that do not generalize well to the broader DHH community.
These issues, according to the authors, limit the real-world applicability of current sign-language technologies.
The Sign-Language Datasheet
To address these gaps, the survey proposes a 24-field Sign-Language Datasheet. This structured template covers aspects such as dataset purpose, sign language variety, number of signers, annotation types, and licensing. The datasheet aims to provide a unified foundation for developers and researchers to evaluate and select datasets more systematically.
Implications for Enterprise AI
For enterprise technology leaders building inclusive AI systems, this survey highlights the importance of dataset quality and standardization. Fragmented data increases development costs and reduces model reliability. The GitHub repository and datasheet offer practical tools for teams evaluating sign-language capabilities in customer service, accessibility, and communication platforms. By adopting standardized documentation, enterprises can reduce annotation inconsistencies and improve model scalability.
As the survey notes, without addressing annotation granularity and signer bias, AI systems risk underperforming in real-world deployments. The 120-dataset index provides a starting point for organizations seeking to benchmark their models against diverse sign languages.
(Note: The original abstract mentions a 24-field Sign-Language Datasheet and a public GitHub repository; this article uses only those facts.)