RÍM — sources and modifications (dictionary version 3) 1. Database of Modern Icelandic Inflection (DMII / BÍN), release 2020.06 Editor: Kristín Bjarnadóttir. Publisher: The Árni Magnússon Institute for Icelandic Studies. Source record: https://repository.clarin.is/repository/xmlui/handle/20.500.12537/5 License: Creative Commons Attribution-ShareAlike 4.0 International. License text: https://creativecommons.org/licenses/by-sa/4.0/legalcode Archive: DIM_2020.06.zip Archive MD5: 0a424a471282290291cb2726a7519da9 Input file: DIM_2020.06_ordmyndir.txt Input SHA-256: d9f3a9d0d773cf9447b8f0762a94f8eeb7751c13ddf3370dc3a33dc916d8573f This app redistributes an adapted word-form index under CC BY-SA 4.0. Changes: NFC normalization; case-insensitive deduplication (prefer a lowercase lexical form where available); removal of entries containing spaces, digits, or punctuation other than internal hyphens; grouping by automatic rhyme key; division into JSON lookup shards. No word forms have been invented or generated. The index contains 3,070,299 distinct normalized word forms, including inflected forms, proper names and uncommon words. It is not a dictionary of definitions or a curated list of 3 million independent headwords. Manifest: ./rhymes-v3/manifest.json Derived data: ./rhymes-v3/1/*.json and ./rhymes-v3/2/*.json These derived data are licensed under CC BY-SA 4.0. Rebuild script in the Site source: scripts/build-dictionary.mjs. 2. Icelandic rhyme algorithms / islenska-org/icelandic Author: Borgar Þorsteinsson (2016). Source: https://github.com/islenska-org/icelandic Revision: 7c778dbf5c0abdb9d75185b47ec2e94c7561192a License: MIT (see ICELANDIC-MIT.txt). Adapted in ../icelandic-rules.mjs. Changes: native ES modules; global replacements; Unicode NFC normalization; vowel-nucleus-based rhyme tails, including short words; hyphen handling; deterministic indexing. Adapted code retains the MIT license. The live islenska.org database was not copied. Its author's openly licensed algorithm is combined with the independent open BÍN data described above. 3. Icelandic Pronunciation Dictionary for Language Technology Authors: Anna Björk Nikulásdóttir, Bjarki Ármannsson, Bryndís Bergþórsdóttir and Eiríkur Rögnvaldsson. Publisher: Grammatek. Source: https://github.com/grammatek/iceprondict Input: ice_pron_dict_standard_clear_IS.csv License: Creative Commons Attribution 4.0 International (CC-BY-4.0.txt). Adaptation: extracted word/SAMPA transcription pairs into ice-pronunciation.tsv. The app compares endings from the last requested vowel nuclei. Vowel length marks are ignored to allow stress-related quantity variation in compounds; vowel quality, consonants and preaspiration are retained. The spelling rules are checked against these transcriptions when both words have pronunciations. Entries whose spoken and written syllable counts disagree are omitted from the phonetic index to avoid spreading the expanded pronunciations of abbreviations to unrelated words. Untranscribed letter-by-letter acronym compounds are omitted from results because written vowels alone do not give their syllable count. When a transcription is absent, spelling rules provide an approximate match; an ending's phonetic evidence can be reused only when all included transcriptions for that written rhyme key agree. Stress is not inferred for the whole compound. Syllable grouping uses pronunciation data when present; otherwise written vowel nuclei are counted, treating au, ei and ey as diphthongs. Unusual spellings, abbreviations and compound boundaries may need human judgment. Rhyme matching does not fully model stress, dialect differences, or every Icelandic sound rule. The original data authors have not reviewed or endorsed the app's rhyme results. 4. Íslensk rímorðabók, Eiríkur Rögnvaldsson, Reykjavík, 1989 Source: user-provided Rim (1).pdf, 831 pages. SHA-256: 13b0ad2261e7219ae04cb0e24ea4ef01d5c377cb53a4f7e052c1e5832edd55cc Derived index: reference-rhymes-v1.json Import script in the Site source: scripts/import-reference.py This separate reference index is a transformation of the supplied PDF, not part of the openly licensed BÍN data. No open-content license is asserted for it. Attribution and original rights are retained; the full PDF is not redistributed. The first whole-word list (PDF pages 16–559) supplies 2,535 rhyme groups containing 19,482 distinct complete word forms. Bold type distinguishes group headings from entries; cross-reference-only headings are not treated as groups. Bound forms prefixed with + or - and other annotations are excluded. The stem-and-suffix tables in the second half are not expanded into invented words. Matching book groups take priority over automatic spelling mergers. An independently listed ending can extend matching to a compound, but its full-word stress is still not inferred. Direct word forms in the matching book group are marked and sorted first within each syllable-count group. The book separates some endings that length-normalized SAMPA merges (e.g. -ás and -áss); those suggestions are placed among looser rhymes. Other book examples checked include hefndi/lemdi and bæinn/daginn, as well as the distinction between falla and fala. ORÐSKÝRINGAR — Íslensk nútímamálsorðabók 2023-12 Ritstjórar: Halldóra Jónsdóttir og Þórdís Úlfarsdóttir. Útgefandi: Stofnun Árna Magnússonar í íslenskum fræðum. Safn: https://repository.clarin.is/repository/xmlui/handle/20.500.12537/318 Leyfi: CC BY-SA 4.0, https://creativecommons.org/licenses/by-sa/4.0/ Afleiddar leitarskrár: data/definitions-2023/*.json, einnig CC BY-SA 4.0. Breytingar: XML fært í 128 JSON-leitarskrár; grunnmyndir normaliseraðar til leitar; valdir málfræðilegir merkimiðar birtir á íslensku. Skýringar, setningardæmi, orðasambönd, svið og notkunarmerki varðveitt með merkingarsamhengi. Árnastofnun ber ekki ábyrgð á framsetningu gagnanna í RÍM. Upprunafærslur: https://islenskordabok.arnastofnun.is/leit/{uppflettiorð} BÍN BEYGINGARTENGINGAR — BinPackage 1.5.0 Source: https://github.com/mideind/BinPackage BÍN editor: Kristín Bjarnadóttir, Árni Magnússon Institute. BÍN data and derived morphology shards: CC BY-SA 4.0. Package software: MIT. Only derived data is redistributed here. Adaptation: lookup IDs for matching ÍNO headwords and word classes; retain real BÍN forms; NFC/lowercase lookup; deduplicate lemma/category pairs; 256 shards. 675,349 keys; restricted to lemmas having ÍNO definitions, not all BÍN entries. ÍSLENSKT ORÐANET — 21.06 Authors: Jón Hilmar Jónsson, Hjalti Daníelsson, Þórður Arnar Árnason, Alec Shaw. Source: https://repository.clarin.is/repository/xmlui/handle/20.500.12537/117 License: CC BY 4.0. Derived data/wordweb shards: CC BY 4.0. Adaptation: extract concept labels and related, broader, narrower relations; merge relations by normalized label; omit missing labels and self-relations. 285,490 concepts; 76,960 lookup keys. Not a list of interchangeable synonyms. 1000 IDIOMATIC EXPRESSIONS — ISLEX, 22.09 Authors: Björn Halldórsson, Árni Davíð Magnússon, Finnur Ágúst Ingimundarson, Einar Freyr Sigurðsson, Steinþór Steingrímsson, Halldóra Jónsdóttir, Þórdís Úlfarsdóttir. Publisher: Árni Magnússon Institute. Source: https://repository.clarin.is/repository/xmlui/handle/20.500.12537/275 License: CC BY 4.0. Derived data/idioms shards: CC BY 4.0. Adaptation: retain Icelandic idiom, sense and both examples from TSV; index by provided keywords; 1000 entries, 779 keyword keys, 256 shards. GREYNIR AND NEFNIR — WORKSHOP Greynir: Miðeind, https://greynir.is/apidoc The selected line is sent on explicit button click to the live IFD tagging API. Greynir is not hosted or downloaded into this static site. Nefnir: Jón Friðrik Daðason and Hrafn Loftsson. Paper: https://aclanthology.org/W19-6133/ Source: https://github.com/jonfd/nefnir Python distribution: nefnir 1.0.2, Sverrir Á. Berg (2019), Apache 2.0. Rules, tags and license retained in data/nefnir. JavaScript port of rule lookup, longest-suffix replacement and recasing in language-resources.mjs. Original license: data/nefnir/LICENSE.txt. Adapted code also Apache 2.0. The analysis is advisory and may be inaccurate for unconventional writing. Rebuild new resource shards: scripts/import-language-resources.py. WIKTIONARY ENGLISH-GLOSS SUPPLEMENT — downloaded 2026-10-11 Authors: Wiktionary contributors. Original pages and author histories: https://en.wiktionary.org/wiki/{word}#Icelandic Machine-readable extraction by Tatu Ylonen / Wiktextract: https://kaikki.org/dictionary/Icelandic/ Input: https://kaikki.org/dictionary/Icelandic/kaikki.org-dictionary-Icelandic.jsonl License selected for text: CC BY-SA 4.0. https://en.wiktionary.org/wiki/Wiktionary:Copyrights https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: retain Icelandic headwords, English glosses and usage tags; omit form-of-only senses, external quotation examples, images and audio. Normalize lookup keys and split into 256 JSON shards. Original headword spelling retained. These adapted glosses remain CC BY-SA 4.0. Each displayed entry links to the original Wiktionary page, whose history credits its contributors. ÍNO definitions take precedence. English fallback is explicitly labelled. No machine translation or invented usage sentences are included. COMPOUND HINTS — BÍN / BinPackage 1.5.0, CC BY-SA 4.0 Rebuild: scripts/import-extra-definitions.py with Wiktionary JSONL and the current BinPackage noun headword list emitted by scripts/current-bin-lemmas.py. Only registered noun compounds with no whole-word dictionary entry are indexed. BÍN's automatic split must match the registered whole lemma and word class; both components must resolve to actual dictionary meanings in ÍNO or Wiktionary. The suffix's noun gender/class must agree with the whole-word entry. BÍN supplies real surface forms, never invented inflections. Additional Wiktionary headwords expand the BÍN inflection index too. Component glosses retain their original source license. Split suggestions are possible analyses, not verified whole-word definitions. Exact whole-word meanings always take precedence. Derived data/compounds and data/compound-plans indices: CC BY-SA 4.0. Manifest files document counts and scope. Existing site sharing is unchanged. For words absent from the precomputed compound index, the browser may suggest a longest-known noun prefix plus a suffix having a real noun definition. Prefixes are drawn from the validated BÍN plans in data/compound-prefixes. These inferred splits are separately labelled; no full-word BÍN registration or full-word meaning is claimed. Unknown components are never defined by guess. WHOLE-WORD DEFINITION FILTER Compact membership index in data/defined-words: ÍNO and Wiktionary headwords, plus their already validated BÍN inflections. No compound-only suggestions. Derived index remains CC BY-SA 4.0. Rebuild after a dictionary update using scripts/build-definition-availability.py. English whole-word meanings count. Filtering preserves rhyme order and syllable grouping and recomputes displayed counts. A failed membership download is retriable, never interpreted as absence.