Methods & Metrics

Methods & Metrics

PRM combines established entropy, lexical-diversity, and repetition measures with three composite indexes that expose different dimensions of corpus behavior.

Use EURE to examine entropy and efficient lexical variation, LDI to examine sustained lexical diversity, and RACS to measure complexity after repetition pressure is accounted for.

Use EURE to examine entropy and efficient lexical variation, LDI to examine sustained lexical diversity, and RACS to measure complexity after repetition pressure is accounted for.

Reading the Metric System

Reading the Metric System

EURE, LDI, and RACS are separate composite lenses built from overlapping but non-identical components. Read each independently, then compare agreement, separation, and rank stability across tokenizers, slices, and corpus segments. Agreement shows that a result persists across multiple measurement families, while divergence identifies which component behavior is driving the difference.

EURE, LDI, and RACS are separate composite lenses built from overlapping but non-identical components. Read each independently, then compare agreement, separation, and rank stability across tokenizers, slices, and corpus segments. Agreement shows that a result persists across multiple measurement families, while divergence identifies which component behavior is driving the difference.

Measurement Scope

Measurement Scope

This page documents how PRM results are assembled from distinct metric families and composite indexes. Reported ranks apply to the tested dataset pool under shared preprocessing, tokenizer, slice, and aggregation conditions. The underlying measurements remain available for inspection alongside the composite results.

This page documents how PRM results are assembled from distinct metric families and composite indexes. Reported ranks apply to the tested dataset pool under shared preprocessing, tokenizer, slice, and aggregation conditions. The underlying measurements remain available for inspection alongside the composite results.

Metric Provenance

Metric Provenance

PRM combines standard external measures with PRM-specific composite indexes. Standard measures include Shannon entropy, unique-token ratio, lexical diversity, MTLD, Maas, Yule’s K, MATTR, MSTTR, Rényi entropy, Tsallis entropy, and normalized Lempel-Ziv complexity.


EURE, LDI, and RACS aggregate selected standardized measures into named analytical views. Every composite result remains traceable to its component metrics, dataset version, tokenizer, slice size, corpus segment, and run output.

PRM combines standard external measures with PRM-specific composite indexes. Standard measures include Shannon entropy, unique-token ratio, lexical diversity, MTLD, Maas, Yule’s K, MATTR, MSTTR, Rényi entropy, Tsallis entropy, and normalized Lempel-Ziv complexity.


EURE, LDI, and RACS aggregate selected standardized measures into named analytical views. Every composite result remains traceable to its component metrics, dataset version, tokenizer, slice size, corpus segment, and run output.

Evidence Charts

Evidence Charts

Begin with the metric-system dashboard, then examine separate index leaderboards, component-level profiles, and side-by-side numeric comparisons.

EURE, LDI, and RACS remain distinct throughout the chart sequence so agreement and separation remain visible.

Begin with the metric-system dashboard, then examine separate index leaderboards, component-level profiles, and side-by-side numeric comparisons.

EURE, LDI, and RACS remain distinct throughout the chart sequence so agreement and separation remain visible.

PRM Metric System Dashboard

PRM Metric System Dashboard

Maps EURE, LDI, and RACS to their component measurements, comparison conditions, and reported aggregate outputs.

Maps EURE, LDI, and RACS to their component measurements, comparison conditions, and reported aggregate outputs.

Open full-size chart

EURE / LDI / RACS leaderboard comparison

EURE / LDI / RACS leaderboard comparison

Displays EURE, LDI, and RACS as separate ranked columns so each index remains visible without being collapsed into a single score.

Displays EURE, LDI, and RACS as separate ranked columns so each index remains visible without being collapsed into a single score.

Open full-size chart

Metric profile heatmap

Metric profile heatmap

Shows the component-level metric profile behind each corpus label, making cross-metric strength, variation, and consistency visible.

Shows the component-level metric profile behind each corpus label, making cross-metric strength, variation, and consistency visible.

Open full-size chart

Reference Comparison Boundary

Reference datasets provide fixed analytical coordinates under shared preprocessing, tokenizer, slice, and metric conditions.

Public-domain references are identified by name, while protected comparator identities and source mappings remain within controlled review.