RMR vs TQT

RMR vs TQT

RMR and TQT are paired analytical transforms of the protected OGO master. RMR preserves lyrical structure for structure-aware analysis, while TQT removes non-lyrical characters for tokenizer and slice-based measurement.

RMR and TQT are paired analytical transforms of the protected OGO master. RMR preserves lyrical structure for structure-aware analysis, while TQT removes non-lyrical characters for tokenizer and slice-based measurement.

RMR

Structure-Preserving

Retains punctuation, line, bar, and verse structure, along with machine-readable verse boundaries for structured metric analysis.

TQT

Text-Normalized

Contains lyrical characters only, with punctuation and symbols normalized to support consistent tokenization and statistical measurement.


Separation

Structure vs Normalization

Routes each analysis to the dataset form containing the information required by that metric.

RMR

RMR

RMR is the structure-preserving analytical dataset, not the human-readable review edition. The protected OGO master contains every verse number, verse title, and creation date. RMR retains verse titles only inside line markers so pipeline scripts can identify where each verse begins and ends.


Because RMR preserves punctuation and line, bar, and verse structure, it is used for cosine similarity and structure-dependent analyses, including rhyme, semantics, and syntax.

RMR is the structure-preserving analytical dataset, not the human-readable review edition. The protected OGO master contains every verse number, verse title, and creation date. RMR retains verse titles only inside line markers so pipeline scripts can identify where each verse begins and ends.


Because RMR preserves punctuation and line, bar, and verse structure, it is used for cosine similarity and structure-dependent analyses, including rhyme, semantics, and syntax.

TQT

TQT

TQT removes every non-lyrical character from the dataset. Punctuation marks and symbols are replaced with a single blank space, except apostrophes, which are removed without separating the surrounding letters.

This prevents apostrophe-related token variants from introducing marginal measurement skew.


TQT is then divided into indiscriminate slice sizes, with each metric measured across every tokenizer at every tested slice size.

TQT removes every non-lyrical character from the dataset. Punctuation marks and symbols are replaced with a single blank space, except apostrophes, which are removed without separating the surrounding letters.

This prevents apostrophe-related token variants from introducing marginal measurement skew.


TQT is then divided into indiscriminate slice sizes, with each metric measured across every tokenizer at every tested slice size.

Distinct Analytical Roles

Distinct Analytical Roles

RMR and TQT preserve different analytical properties from the same underlying corpus. RMR retains the structural information required for rhyme, semantic, syntactic, and verse-aware analysis. TQT removes that structure to provide normalized lyrical text for tokenizer comparison and slice-based statistical measurement.


Neither transform replaces the protected OGO master. Each exists to support a specific analytical workload.

RMR and TQT preserve different analytical properties from the same underlying corpus. RMR retains the structural information required for rhyme, semantic, syntactic, and verse-aware analysis. TQT removes that structure to provide normalized lyrical text for tokenizer comparison and slice-based statistical measurement.


Neither transform replaces the protected OGO master. Each exists to support a specific analytical workload.

Pipeline Routing and Split Rules

Pipeline Routing and Split Rules

OGO remains the fully labeled source master. RMR routes structure-dependent analyses, while TQT routes normalized tokenizer and slice-based measurement.


When datasets are divided into H1/H2 or T1/T2/T3 formats, every split occurs only between complete verses. No split cuts through a verse. This preserves rhyme patterns, local semantic relationships, and syntactic structures without introducing content-driven edits.

OGO remains the fully labeled source master. RMR routes structure-dependent analyses, while TQT routes normalized tokenizer and slice-based measurement.


When datasets are divided into H1/H2 or T1/T2/T3 formats, every split occurs only between complete verses. No split cuts through a verse. This preserves rhyme patterns, local semantic relationships, and syntactic structures without introducing content-driven edits.

Dataset Lineage

Dataset Lineage

OGO is the fully labeled source master. RMR is the structure-preserving analytical layer. TQT is the normalized lyrical-text layer. All three represent the same underlying corpus, remain versioned together, and are routed to analyses based on the information each version preserves.

OGO is the fully labeled source master. RMR is the structure-preserving analytical layer. TQT is the normalized lyrical-text layer. All three represent the same underlying corpus, remain versioned together, and are routed to analyses based on the information each version preserves.

How to read the charts

How to read the charts

Begin with deep metric coverage to see the analytical surface downstream of RMR and TQT. Tokenizer robustness shows whether TQT-based results persist across tokenization systems. Raw trends and the metric-slice heatmap show how results behave across measurement windows, while the 65536 macro lens examines corpus-scale behavior.

Structure-dependent charts should be read as RMR outputs because those analyses rely on retained punctuation and line, bar, and verse boundaries.

Begin with deep metric coverage to see the analytical surface downstream of RMR and TQT. Tokenizer robustness shows whether TQT-based results persist across tokenization systems. Raw trends and the metric-slice heatmap show how results behave across measurement windows, while the 65536 macro lens examines corpus-scale behavior.

Structure-dependent charts should be read as RMR outputs because those analyses rely on retained punctuation and line, bar, and verse boundaries.

Protected Dataset Materials

Protected Dataset Materials

This page describes transform roles and aggregate outputs. It does not publish raw writing, protected excerpts, private transform files, source-level mappings, third-party source labels, artist names, album titles, or song titles.

This page describes transform roles and aggregate outputs. It does not publish raw writing, protected excerpts, private transform files, source-level mappings, third-party source labels, artist names, album titles, or song titles.

Disclosure Boundary

Disclosure Boundary

Public pages disclose dataset roles, transformation rules, aggregate results, metric behavior, and method provenance. Protected corpus files, identities, titles, and reconstructable mappings remain within controlled review.

Public pages disclose dataset roles, transformation rules, aggregate results, metric behavior, and method provenance. Protected corpus files, identities, titles, and reconstructable mappings remain within controlled review.

Project Rocket Man