iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million
Home ›› Technology ›› Ai ›› Toten Framework Outperforms Statistical Tokenization for Physical Quantities in Brazilian Portuguese Technical Texts

Toten Framework Outperforms Statistical Tokenization for Physical Quantities in Brazilian Portuguese Technical Texts

Researchers present Toten, a framework that replaces statistical tokenization with ontology-based classification for physical quantities and technical notation in Brazilian Portuguese. The system leverages external oracles and achieves higher ontological atomicity and numerical reconstruction compared to state-of-the-art baselines.

iG
iGEN Editorial
July 8, 2026
Toten Framework Outperforms Statistical Tokenization for Physical Quantities in Brazilian Portuguese Technical Texts

Statistical tokenization methods like Byte-Pair Encoding (BPE) are efficient for vocabulary compression but semantically blind to structured technical entities, often fragmenting physical quantities, numbers, units, and symbolic expressions into arbitrary subwords. According to a paper published on arXiv, researchers have introduced Toten, a knowledge-based ontological tokenization framework that replaces statistical derivation with declarative classification grounded in a formal ontology of engineering entities (OEE).

What Is Toten?

Toten is formalized as the triple <O, classify, {inst_tau}>: the ontology gathers types, structural principles, composition relations, and preservable invariants; the classification function maps raw text into typed regions; and the instantiator family yields a self-descriptive structured representation. The framework’s robustness derives from deterministic coupling with three external oracles: Pint (dimensional analysis), Unicode Character Database (typographic properties), and RSLP (Portuguese morphology). The authors are Antonio de Sousa Leitão Filho, Allan Kardec Duailibe Barros Filho, Fabrício Saul Lima, Selby Mykael Lima dos Santos, and Rejani Bandeira Vieira.

Evaluation Against Baselines

The paper reports intrinsic evaluation covering four properties verifiable by construction: ontological atomicity, dimensional equivalence, typographic robustness, and numerical reconstruction. Testing was conducted on an internal benchmark (EngQuant, N=800) and four Brazilian Portuguese external corpora (N=1771 eligible cases). Against eight state-of-the-art baselines, Toten achieved unit ontological atomicity in all contrasts. Numerical reconstruction scores on external corpora ranged from 0.775 to 0.904, compared to 0.627–0.703 for the best baseline, Quantulum3. On EngQuant, Toten scored 0.780 versus 0.340 for Quantulum3. The authors report that differences are statistically significant (McNemar with Holm correction). Spearman correlation between internal and external rankings confirmed concurrent validity of the control benchmark. Dimensional equivalence showed statistical parity with Pint, the oracle from which the system inherits dimensional authority.

Metric Toten Best Baseline (Quantulum3)
Numerical reconstruction (external corpora) 0.775–0.904 0.627–0.703
Numerical reconstruction (EngQuant) 0.780 0.340
Ontological atomicity Unit (all contrasts) Not achieved

Implications for Technical Document Processing

While the paper does not explicitly address commercial applications, the ability to accurately tokenize physical quantities and technical notation in Brazilian Portuguese has direct relevance for any system processing engineering, scientific, or technical documentation. Enterprise technology leaders dealing with multilingual technical content may find Toten’s ontology-driven approach a significant improvement over general-purpose tokenizers. The framework’s deterministic coupling with well-established oracles ensures consistency and dimensional correctness, which is critical for automated document understanding in industries like manufacturing, energy, and logistics.


Sources:

Keep Reading

Recommended Stories

Large Language Models Can Read Compressed Text That Humans Cannot, Researchers Find Technology

Large Language Models Can Read Compressed Text That Humans Cannot, Researchers Find

A new research paper introduces BabelTele, a compact, non-human-readable text format that large language models can still interpret with high semantic fidelity. The approach compresses text to 27.9% of its original length while preserving 99.5% of meaning, potentially reducing context overhead and costs in enterprise AI deployments.

June 20, 2026
G2Rec Framework Structures and Tokenizes User Interests for Generative Recommendation Technology

G2Rec Framework Structures and Tokenizes User Interests for Generative Recommendation

The G2Rec framework, proposed by researchers, addresses limitations in generative recommendation by unifying holistic graph-based user co-engagement modeling with semantic tokenization. It enables scalable, accurate user interest modeling without requiring ground-truth interests, and has demonstrated superiority through online deployment and experiments on public datasets.

June 20, 2026
Koshur Diacritizer: A Byte-Level Model Restores Diacritics for Kashmiri Language NLP Technology

Koshur Diacritizer: A Byte-Level Model Restores Diacritics for Kashmiri Language NLP

Researchers have developed Koshur Diacritizer, a byte-level sequence-to-sequence model based on ByT5-small, to restore missing diacritic marks in Kashmiri digital text. The model, trained on 23,700 sentence pairs, achieves a DERm of 0.2012 and word error rate of 0.2159, with a native expert accuracy of 77.5%. The dataset, model, and source code are publicly released to support low-resource language research.

June 16, 2026
Creating Multilingual Mental Health Datasets: Study Reveals Limits of Persona-Based Localization via Nationality and Language Technology

Creating Multilingual Mental Health Datasets: Study Reveals Limits of Persona-Based Localization via Nationality and Language

A new arxiv paper investigates whether persona-based methods can generate multilingual mental health dialogue datasets by modifying nationality and language. The study found that just adding these parameters introduces clinical inconsistencies across languages, and LLM judge models exhibit inaccuracies in assessing depression severity in non-English texts, highlighting the need for culturally responsive data generation.

July 8, 2026