Hayzunxn
A joke language kept like software: a TSV lexicon, a versioned spec and 12 lints
- Role
- Solo, directing agents
- Status
- Shipped
- Source
- Public on GitHub
- Stack
- PythonMarkdownTSVAutoHotkey
In numbers
473
lexicon entries
12
validation lints, all passing at v2.0.0
4
documents regenerated from the lexicon and spec: a Word grammar, a spreadsheet dictionary, an overview deck and a PDF
The problem
Hayzûnxñ began as a joke language I made with a friend before I wrote any code, and grew later with AI help. It is meant to sound funny: the register is crude on purpose, homonym collisions are intentional, and some of the phonology is unpronounceable by design. A language built on deliberate 'bugs' needs a way to tell a planned oddity from a real mistake, and loose documents could not.
The approach
I later had agents restructure it the way a linguist would and treat it like code. The dictionary became data, the grammar became numbered spec modules, every change runs through a validator, and every document is generated from the two.
How it works
The lexicon is the database
lexicon.tsv holds 473 entries, each with a stable ID such as N0001 or V0001, and the spec refers to words by ID instead of by spelling, so a respelled word breaks nothing. Every entry carries its origin (original core, audited expansion, or v2 coinage) and its register, so nothing is rewritten silently.
Lints that know a joke from a bug
validate.py runs 12 lints, among them the ñ ending rule, undefined letters, broken cross-references and homonyms. A homonym fails unless it is marked as intentional, so Yukñ can mean both south and how on purpose while an accidental clash stops the build.
Decisions locked in writing
A rationale module lists the choices that look like bugs and are not. The future prefix ñ- cannot be pronounced at the start of a syllable, and that is a feature. The contributing guide sends anyone opening a pull request to that module first.
Two scripts and generated documents
Latin script is primary and Cyrillic secondary, with a transliterator between them and an AutoHotkey script for typing the letters. build.py regenerates the grammar as a Word document, the dictionary as a spreadsheet, an overview deck and a PDF of the spec. The repo also ships an olympiad-style puzzle and a bootstrap document that teaches the language to a language model.
What I chose, and what lost
Chose
Markdown for prose and TSV for data, with the Office files generated
Over
Editing the grammar and the dictionary as Word and Excel files
Both text formats diff cleanly in git, and a generated file cannot drift from its source.
Chose
Dual licence: MIT for the code, CC BY-SA 4.0 for the language
Over
One licence for the whole repository
Code and content sit on either side of a natural seam, and share-alike keeps derivative dictionaries and forks of the language open.
Outcome
Version 2.0.0 shipped with all 12 lints passing, and the repository went public after a legacy folder was scrubbed from its git history. The README calls the crude register core, so this page quotes the clean words. Open for v3: a translated corpus, a mythology, about 500 non-Bulgarian roots and one unusual typological feature. Later I considered turning it into an esoteric programming language, hard for a person to write and easy for a model, and dropped the idea.