RMR vs TQT
RMR vs TQT
RMR and TQT are paired analytical transforms of the protected OGO master. RMR preserves lyrical structure for structure-aware analysis, while TQT removes non-lyrical characters for tokenizer and slice-based measurement.
RMR and TQT are paired analytical transforms of the protected OGO master. RMR preserves lyrical structure for structure-aware analysis, while TQT removes non-lyrical characters for tokenizer and slice-based measurement.
RMR
Structure-Preserving
Retains punctuation, line, bar, and verse structure, along with machine-readable verse boundaries for structured metric analysis.
TQT
Text-Normalized
Contains lyrical characters only, with punctuation and symbols normalized to support consistent tokenization and statistical measurement.
Separation
Structure vs Normalization
Routes each analysis to the dataset form containing the information required by that metric.
RMR
RMR
RMR is the structure-preserving analytical dataset, not the human-readable review edition. The protected OGO master contains every verse number, verse title, and creation date. RMR retains verse titles only inside line markers so pipeline scripts can identify where each verse begins and ends.
Because RMR preserves punctuation and line, bar, and verse structure, it is used for cosine similarity and structure-dependent analyses, including rhyme, semantics, and syntax.
RMR is the structure-preserving analytical dataset, not the human-readable review edition. The protected OGO master contains every verse number, verse title, and creation date. RMR retains verse titles only inside line markers so pipeline scripts can identify where each verse begins and ends.
Because RMR preserves punctuation and line, bar, and verse structure, it is used for cosine similarity and structure-dependent analyses, including rhyme, semantics, and syntax.
TQT
TQT
TQT removes every non-lyrical character from the dataset. Punctuation marks and symbols are replaced with a single blank space, except apostrophes, which are removed without separating the surrounding letters.
This prevents apostrophe-related token variants from introducing marginal measurement skew.
TQT is then divided into indiscriminate slice sizes, with each metric measured across every tokenizer at every tested slice size.
TQT removes every non-lyrical character from the dataset. Punctuation marks and symbols are replaced with a single blank space, except apostrophes, which are removed without separating the surrounding letters.
This prevents apostrophe-related token variants from introducing marginal measurement skew.
TQT is then divided into indiscriminate slice sizes, with each metric measured across every tokenizer at every tested slice size.
Distinct Analytical Roles
Distinct Analytical Roles
RMR and TQT preserve different analytical properties from the same underlying corpus. RMR retains the structural information required for rhyme, semantic, syntactic, and verse-aware analysis. TQT removes that structure to provide normalized lyrical text for tokenizer comparison and slice-based statistical measurement.
Neither transform replaces the protected OGO master. Each exists to support a specific analytical workload.
RMR and TQT preserve different analytical properties from the same underlying corpus. RMR retains the structural information required for rhyme, semantic, syntactic, and verse-aware analysis. TQT removes that structure to provide normalized lyrical text for tokenizer comparison and slice-based statistical measurement.
Neither transform replaces the protected OGO master. Each exists to support a specific analytical workload.
Pipeline Routing and Split Rules
Pipeline Routing and Split Rules
OGO remains the fully labeled source master. RMR routes structure-dependent analyses, while TQT routes normalized tokenizer and slice-based measurement.
When datasets are divided into H1/H2 or T1/T2/T3 formats, every split occurs only between complete verses. No split cuts through a verse. This preserves rhyme patterns, local semantic relationships, and syntactic structures without introducing content-driven edits.
OGO remains the fully labeled source master. RMR routes structure-dependent analyses, while TQT routes normalized tokenizer and slice-based measurement.
When datasets are divided into H1/H2 or T1/T2/T3 formats, every split occurs only between complete verses. No split cuts through a verse. This preserves rhyme patterns, local semantic relationships, and syntactic structures without introducing content-driven edits.
Dataset Lineage
Dataset Lineage
OGO is the fully labeled source master. RMR is the structure-preserving analytical layer. TQT is the normalized lyrical-text layer. All three represent the same underlying corpus, remain versioned together, and are routed to analyses based on the information each version preserves.
OGO is the fully labeled source master. RMR is the structure-preserving analytical layer. TQT is the normalized lyrical-text layer. All three represent the same underlying corpus, remain versioned together, and are routed to analyses based on the information each version preserves.
How to read the charts
How to read the charts
Begin with deep metric coverage to see the analytical surface downstream of RMR and TQT. Tokenizer robustness shows whether TQT-based results persist across tokenization systems. Raw trends and the metric-slice heatmap show how results behave across measurement windows, while the 65536 macro lens examines corpus-scale behavior.
Structure-dependent charts should be read as RMR outputs because those analyses rely on retained punctuation and line, bar, and verse boundaries.
Begin with deep metric coverage to see the analytical surface downstream of RMR and TQT. Tokenizer robustness shows whether TQT-based results persist across tokenization systems. Raw trends and the metric-slice heatmap show how results behave across measurement windows, while the 65536 macro lens examines corpus-scale behavior.
Structure-dependent charts should be read as RMR outputs because those analyses rely on retained punctuation and line, bar, and verse boundaries.
Protected Dataset Materials
Protected Dataset Materials
This page describes transform roles and aggregate outputs. It does not publish raw writing, protected excerpts, private transform files, source-level mappings, third-party source labels, artist names, album titles, or song titles.
This page describes transform roles and aggregate outputs. It does not publish raw writing, protected excerpts, private transform files, source-level mappings, third-party source labels, artist names, album titles, or song titles.
Disclosure Boundary
Disclosure Boundary
Public pages disclose dataset roles, transformation rules, aggregate results, metric behavior, and method provenance. Protected corpus files, identities, titles, and reconstructable mappings remain within controlled review.
Public pages disclose dataset roles, transformation rules, aggregate results, metric behavior, and method provenance. Protected corpus files, identities, titles, and reconstructable mappings remain within controlled review.
Related
Project Rocket Man
Project Rocket Man
