iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb Benchmark Diesel Price Rises Third Week as Futures Plunge; Spread Hits Record Indian Government Limits Sugar Dealers to 400 Tonnes Stock Until November to Curb Hoarding Tenants signing longer leases for larger warehouses as 3PLs lock in capacity US stock market flat as S&P 500 and Dow barely move, Nasdaq slides over 1% on chip rout Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb Benchmark Diesel Price Rises Third Week as Futures Plunge; Spread Hits Record Indian Government Limits Sugar Dealers to 400 Tonnes Stock Until November to Curb Hoarding Tenants signing longer leases for larger warehouses as 3PLs lock in capacity US stock market flat as S&P 500 and Dow barely move, Nasdaq slides over 1% on chip rout
Home ›› Technology ›› Ai ›› Comprehensive Survey of 120 Sign-Language Datasets Identifies Key Gaps in Scale and Annotation Standards

Comprehensive Survey of 120 Sign-Language Datasets Identifies Key Gaps in Scale and Annotation Standards

A comprehensive survey of 120 sign-language datasets across 35 languages reveals fragmented annotations, modality imbalance, and signer bias. The study introduces a 24-field datasheet and a public GitHub repository to standardize documentation and improve scalability.

iG
iGEN Editorial
June 20, 2026
Comprehensive Survey of 120 Sign-Language Datasets Identifies Key Gaps in Scale and Annotation Standards

Sign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities, yet advances in sign-language recognition, translation, and production remain constrained by fragmented datasets, inconsistent annotations, and limited linguistic coverage. A new survey, published on arXiv, provides a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages, and systematically analyzes key challenges such as modality imbalance, annotation granularity, and signer bias.

According to the survey by researchers Ni, Yiming; Cheng, Zhi-Qi; Li, Jiayu; and Wei, existing benchmarks often fail to reflect real-world communication needs. The work outlines considerations for future dataset design and introduces a 24-field Sign-Language Datasheet to support standardized documentation. The authors also released a public GitHub repository to enable reproducible evaluation.

Survey Scope and Scale

The study catalogs 120 distinct sign-language datasets spanning 35 different sign languages. This breadth highlights the diversity of resources available but also underscores fragmentation. The researchers note that while some languages like American Sign Language (ASL) and Chinese Sign Language (CSL) are well-represented, many others have few or no large-scale datasets.

Metric Value
Number of datasets 120
Sign languages covered 35
Datasheet fields 24
Repository Public GitHub

Key Challenges Identified

The survey identifies three primary challenges:

  • Modality Imbalance: Most datasets rely on single camera views or limited sensor modalities, failing to capture the full expressiveness of signing.
  • Annotation Granularity: Annotations vary widely — from coarse gloss labels to precise linguistic features — making cross-dataset comparisons difficult.
  • Signer Bias: Many datasets include a small number of signers, leading to models that do not generalize well to the broader DHH community.

These issues, according to the authors, limit the real-world applicability of current sign-language technologies.

The Sign-Language Datasheet

To address these gaps, the survey proposes a 24-field Sign-Language Datasheet. This structured template covers aspects such as dataset purpose, sign language variety, number of signers, annotation types, and licensing. The datasheet aims to provide a unified foundation for developers and researchers to evaluate and select datasets more systematically.

Implications for Enterprise AI

For enterprise technology leaders building inclusive AI systems, this survey highlights the importance of dataset quality and standardization. Fragmented data increases development costs and reduces model reliability. The GitHub repository and datasheet offer practical tools for teams evaluating sign-language capabilities in customer service, accessibility, and communication platforms. By adopting standardized documentation, enterprises can reduce annotation inconsistencies and improve model scalability.

As the survey notes, without addressing annotation granularity and signer bias, AI systems risk underperforming in real-world deployments. The 120-dataset index provides a starting point for organizations seeking to benchmark their models against diverse sign languages.

(Note: The original abstract mentions a 24-field Sign-Language Datasheet and a public GitHub repository; this article uses only those facts.)


Sources:

Keep Reading

Recommended Stories

PrefSQA Introduces Pairwise Preference Prediction for Speech Quality Assessment Technology

PrefSQA Introduces Pairwise Preference Prediction for Speech Quality Assessment

A research paper proposes PrefSQA, a pairwise preference prediction method for speech quality assessment that reduces rater variability compared to traditional mean opinion scores. The method incorporates uncertainty-aware logits, an impairment attention head, and non-matching-reference comparisons. Experiments on five datasets show clear improvements over baselines, especially with high-quality preference data.

July 8, 2026
Unsupervised Algorithms Cut Annotation Time by 78% for Industrial Semantic Segmentation Technology

Unsupervised Algorithms Cut Annotation Time by 78% for Industrial Semantic Segmentation

Researchers have demonstrated that unsupervised computer vision algorithms can reduce the annotation time for semantic segmentation tasks in industrial materials science by 78%, from 170 hours to 37 hours. The team created the largest public steel microstructure segmentation dataset and a benchmark deep learning model, validated by field experts and deployed in an industrial setting.

June 21, 2026
New Survey Maps How Evidence Tracing and Execution Provenance Can Make LLM Agents Trustworthy Technology

New Survey Maps How Evidence Tracing and Execution Provenance Can Make LLM Agents Trustworthy

A new survey from arXiv explores evidence tracing and execution provenance as key mechanisms for ensuring trustworthiness in LLM-based agents. The paper defines a unified framework connecting retrieval grounding, tool-use safety, memory lineage, and failure diagnosis, and reviews benchmarks and open challenges.

June 16, 2026
Medical Image Segmentation Survey: U-Net, Transformers, SAM and Clinical Translation Challenges Technology

Medical Image Segmentation Survey: U-Net, Transformers, SAM and Clinical Translation Challenges

A new arXiv survey systematically reviews medical image segmentation methods based on U-Net, Transformer, and SAM architectures. It covers public datasets, evaluation metrics, and key challenges, aiming to guide future research and clinical adoption. The authors have made all related resources publicly available on GitHub.

June 16, 2026