Skip to the content
Georgi DimitrovdaTuzzo

Hayzunxn

A joke language kept like software: a TSV lexicon, a versioned spec and 12 lints

Role
Solo, directing agents
Status
Shipped
Source
Public on GitHub
Stack
PythonMarkdownTSVAutoHotkey
ñññññññññññññññ

In numbers

473

lexicon entries

12

validation lints, all passing at v2.0.0

4

documents regenerated from the lexicon and spec: a Word grammar, a spreadsheet dictionary, an overview deck and a PDF

The problem

Hayzûnxñ began as a joke language I made with a friend before I wrote any code, and grew later with AI help. It is meant to sound funny: the register is crude on purpose, homonym collisions are intentional, and some of the phonology is unpronounceable by design. A language built on deliberate 'bugs' needs a way to tell a planned oddity from a real mistake, and loose documents could not.

The approach

I later had agents restructure it the way a linguist would and treat it like code. The dictionary became data, the grammar became numbered spec modules, every change runs through a validator, and every document is generated from the two.

How it works

  1. The lexicon is the database

    lexicon.tsv holds 473 entries, each with a stable ID such as N0001 or V0001, and the spec refers to words by ID instead of by spelling, so a respelled word breaks nothing. Every entry carries its origin (original core, audited expansion, or v2 coinage) and its register, so nothing is rewritten silently.

  2. Lints that know a joke from a bug

    validate.py runs 12 lints, among them the ñ ending rule, undefined letters, broken cross-references and homonyms. A homonym fails unless it is marked as intentional, so Yukñ can mean both south and how on purpose while an accidental clash stops the build.

  3. Decisions locked in writing

    A rationale module lists the choices that look like bugs and are not. The future prefix ñ- cannot be pronounced at the start of a syllable, and that is a feature. The contributing guide sends anyone opening a pull request to that module first.

  4. Two scripts and generated documents

    Latin script is primary and Cyrillic secondary, with a transliterator between them and an AutoHotkey script for typing the letters. build.py regenerates the grammar as a Word document, the dictionary as a spreadsheet, an overview deck and a PDF of the spec. The repo also ships an olympiad-style puzzle and a bootstrap document that teaches the language to a language model.

What I chose, and what lost

Chose

Markdown for prose and TSV for data, with the Office files generated

Over

Editing the grammar and the dictionary as Word and Excel files

Both text formats diff cleanly in git, and a generated file cannot drift from its source.

Chose

Dual licence: MIT for the code, CC BY-SA 4.0 for the language

Over

One licence for the whole repository

Code and content sit on either side of a natural seam, and share-alike keeps derivative dictionaries and forks of the language open.

Outcome

Version 2.0.0 shipped with all 12 lints passing, and the repository went public after a legacy folder was scrubbed from its git history. The README calls the crude register core, so this page quotes the clean words. Open for v3: a translated corpus, a mythology, about 500 non-Bulgarian roots and one unusual typological feature. Later I considered turning it into an esoteric programming language, hard for a person to write and easy for a model, and dropped the idea.