Status and limitations¶
v0.4.0-alpha. omnist-j implements the full Document model, Schema model, OML and OSD
grammars (read and write), validate, materialize, the full schema
algebra (satisfiable_set, is_empty, prune, compatible_with,
equivalent, normalize, extract, lint, infer), all four
interchange codecs (JSON/YAML/TOML/XML, read and write), and a CLI.
Conformance¶
Both tracks of the conformance harness run against vendor/omnist-spec
v0.28.0-beta's pinned suite (commit 1a7d0de), comparing diagnostics as (path, code) sets
(omnist-spec §8.5.2 rules 1-3: code-aware, not code-agnostic; no path or code
loosening anywhere in the runner). 0 failures.
| Track | Pass | Fail | Skip |
|---|---|---|---|
| 1: OML/OSD CLI fixtures | 29 (19 comparable*) | 0 | 0 |
| 2: JSON vectors | 310 | 0 | 28 |
* The Java harness folds the 10 _referee-self-test/* fixtures into its Track 1
headline (omnist-j#110); the other ports report 19.
v0.27.0-beta to v0.28.0-beta (2026-10-02). No vector changed: Track 2 stays 310 pass / 0 fail / 28 skip
of 338, Track 1 29 / 0 / 0, headline 339 / 0 / 28. The release adds four rules about programmatically built
schemas that no vector can reach (DIV-5), so they are pinned by SchemaV028Test only; see
Programmatic schema rules below.
v0.26.0-beta to v0.27.0-beta (2026-10-02). Track 2 went from 303 pass / 0 fail / 28 skip (of 331
vectors) to 310 / 0 / 28 (of 338); the harness headline from 332 / 0 / 28 to 339 / 0 / 28. The 7 new vectors pin the
empty merge sequence <<: [] (D-18a), a carrier that merges nothing; they passed with no code change.
v0.22.0-beta to v0.26.0-beta (2026-10-01). Track 2 went from 253 pass / 0 fail / 34 skip (of 287
vectors) to 303 pass / 0 fail / 28 skip (of 331); the harness headline from 282 / 0 / 34 to 332 / 0 / 28
(Track 1 unchanged at 29 / 0 / 0). Measured red to green: the bump alone, with the old runner, gave 264 pass /
8 fail / 59 skip (the 8 are the E-32 line:col placeholder vectors the runner did not yet understand, 4 of them
in YAML's alias-expansion file; 31 vectors carrying declared_max_alias_expansion or declared_max_expanded_slots
were skipped); with the runner honouring both keys and the E-32 placeholder, but D-18 and D-22 not yet
enforced, 287 pass / 16 fail / 28 skip (the 16 are every alias-expansion and expanded-size vector that expects
a refusal); with D-18, D-18a, D-19, D-20 and D-22 enforced, 303 / 0 / 28. The only skips left are the 28
OSD-OML vectors (omnist-j#105). All 35 vectors of formats-yaml/alias-expansion.json run and pass,
each against the number it declares (the runner passes a declared maximum to the reader only for a vector
that carries the key), including the 5 expanded-size ones and the 5 malformed-merge syntax ones; the E-32
placeholder is matched as the spec's E-32b says (same code, a well-formed line:col, nothing closer).
v0.21.0-beta to v0.22.0-beta (2026-09-29). Track 2 went from 241 pass / 12 fail / 34 skip
(of 287 vectors, measured with the bump applied and no code change) to 253 pass / 0 fail / 34 skip;
the harness headline from 270 / 12 / 34 to 282 / 0 / 34 (Track 1 unchanged at 29 / 0 / 0). The 12
failures were DIV-7's: ten OML-26 separator-then-stray-token vectors
(shape/separator-then-closing-brace-..., semicolon-then-closing-brace-...,
separator-then-comma-..., separator-then-closing-bracket-..., separator-then-a-non-label-token-...,
separator-then-nan-..., separator-then-opening-brace-..., separator-then-an-array-...,
separator-then-colon-..., and closing-brace-after-a-braced-edge-value-..., all reported
parse.unexpected-token) and the two E-28 code-point column vectors
(oml-grammar/errors/column-counts-code-points-after-an-astral-character at 1:13 for 1:12,
osd-grammar/errors/... at 2:19 for 2:18). Changes: after a complete top-level edge the edge
list continues only if a separator is followed by a STRING or IDENT, and any other leftover token
is parse.trailing-content with or without a separator (OML-26, OML-25); inside {...} and [...]
it stays parse.unexpected-token (OML-27); the OML and OSD lexers count the column in Unicode code
points (E-28), so the second half of a surrogate pair adds nothing. Skips are unchanged at 34.
Measured on one-line inputs of about 1.9 million characters (the input cap is 2,000,000), before
and after, best of three: 200k-token array 207 ms / 184 ms, 100k edges 347 ms / 285 ms, a 1.9M-character
string 19 ms / 21 ms, an OSD line of 100k fields 98 ms / 73 ms. The column computation stays linear.
v0.19.0-beta to v0.21.0-beta (2026-09-22). Track 2 went from 215 pass / 0 fail / 34 skip
(of 249 vectors) to 239 pass / 0 fail / 34 skip (of 273): 215 + 14 bytes_hex D-14 vectors
(all presented through the CLI's byte-oriented entry point, per E-27; none decoded with
replacement) + 5 OML-26/OML-27 stray-token vectors (4 shape/* + 1 arrays/*, already
satisfied, unmeasured before this sweep) + 4 OSD-15 canonical-escaping vectors (already
satisfied, unmeasured before this sweep) + 1
infer/allow-any/mixed-object-and-scalar-shapes-open-to-any-when-allowed vector (already
satisfied) = 239. Skip count and its breakdown are unchanged at 34 (28 OSD-OML + 6
alias-expansion, DIV-3) -- both the pre-sweep and post-sweep runs give the identical
25 parse_schema_oml + 3 write_schema_oml + 6 alias-expansion split. Fixed a real
conformance-runner gap found during this sweep's comparison audit: runParseSchemaVector
never compared
expect.schema on a successful parse_schema vector at all (not even structurally) --
confirmed by mutation, a corrupted expect.schema on
osd-grammar/canonical-output/declaration-order-round-trips-exactly still reported PASS
before the fix. Now compared byte-for-byte per Sec3.3/Sec5.9, and the same mutation now
fails as expected.
Every skip is a real, cited reason (omnist-spec §8.5.5, E-20), never a run against a wrong default:
| Skips | Reason |
|---|---|
| 28 | Not yet implemented: the OSD-OML extension (omnist-j#105); 25 parse_schema_oml, 3 write_schema_oml |
A vector whose operation the runner has never heard of is a failure, and a failure whose exception
carries no structured code/path is a failure, not a guess. There is no runner-side "known failing"
list: the CI gate fails on any nonzero fail count (E-22).
YAML alias limits (D-18, D-18a, D-19, D-20, D-22)¶
Implemented in 0.3.0-alpha (omnist-j#113),
against omnist-spec v0.26.0-beta (commit 7744a5c). YamlCodec.read(text) composes the text into
SnakeYAML's node graph, which describes anchors and aliases without expanding them, and checks that
graph before anything is constructed from it. The check is one iterative pass in time linear in
the input, with each container's counts memoized, and saturating arithmetic.
Two limits, two codes, both finite and both configurable through YamlLimits (D-10, D-11):
| Limit | Default | Option range | Code | What it bounds |
|---|---|---|---|---|
Expansion factor E = W / S of every mapping and sequence |
50 | 1 to 10 000 | document.limit.alias-expansion |
Amplification: how many value slots a written construct materializes per slot written |
Expanded size W(root) |
1 000 000 | 1 to 10 000 000 | document.limit.expanded-size |
Absolute size of an input that uses anchors |
- What is checked. Every mapping and every sequence in a value position, the document root, an
inline merge source and every anchored definition. Scalars are never checked. A sequence in
merge-value position (
<<: [*a, *b], anchored or not) is a carrier: no slot, not a candidate (D-18a). - The ratio of a merge is not 1.00. A mapping that merges a block of
kkeys and writes one key of its own hasEof about(k + 2) / 3:job: {<<: *base, script: x}writes three slots and materializesk + 2. So 100 services merging a 20-key block read at most 7.33 and 100 merging a 60-key block 20.67, but onejobmerging a 150-key base reads 50.67 and is refused at the default, and so is a 100-key block aliased 100 times at the root (50.50). A workload that needs more raises the maximum (up to 10 000); the previous behaviour, SnakeYAML's global count of 50 aliases to collections, refused any configuration that held more than 50 of them, however harmless. - The two limits are independent. An input can pass the ratio and fail the size (a large document
whose every container sits under 50), or pass the size and fail the ratio (a small, very amplifying
one). An input that fails both is reported as
document.limit.alias-expansion. - Exemption.
document.limit.expanded-sizeapplies only to an input that contains at least one alias or merge key (a<<key, as the spec uses the term; a quoted"<<"is a plain key). A YAML input with neither is treated as a JSON or OML input of the same size is: the 2,000,000-character input cap and the node limit govern it. The cliff is deliberate in the spec: a plain file of two million slots passes and adding one alias subjects it to the cap. In this port the exemption applies regardless, but a plain file that large cannot reach the reader: the 2,000,000-character input cap and the 1,000,000-node limit refuse it first, so in practice no YAML input here is exempt from a cap. Wis conservative. It ignores key collisions, so a document whose merged keys are overridden can be refused though it materializes fewer slots (D-19).- Malformed merges (
<<: 1,<<: [1],<<: [[{a: 1}]],<<: *sover scalars) areparse.codec-syntaxand win over both limit codes, wherever they sit in the document. - Cycles. An anchor that refers to itself, directly, through other anchors or through a merge
key (
a: &a {<<: *a}), isdocument.limit.alias-expansion(D-20). - Parse cost. The check is cheap next to composing the text, which is SnakeYAML's own work and is paid
first, in proportion to the input; for an input near the 2,000,000-character cap composing is the dominant cost
of refusing it. Measured (best of seven, on the composed graph, JDK 21, this repository's test JVM): the
fan-out bombs 4^10, 2^30 and 50^4 (465 to 2 009 characters) compose in 2.7 to 3.8 ms and are checked in 0.2 to
0.5 ms; the unanchored merge fan-in of 2 000 merges of a 2 000-key block (31 803 characters) composes in
15 ms and is checked in 2.7 ms; a root list of 100 000
{k: *b}over 1 000 scalars (1.2 MB) composes in 325 ms and is checked in 56 ms (the merge-shape pass reads the whole document before counting); a legitimate 1 MB file of 40 000 services merging a 20-key block composes in 295 ms, is checked in 50 ms and is accepted, and its full read, which materializes about 880 000 slots, takes 0.9 s. - Interaction with SnakeYAML's own limits. The library's global cap (
maxAliasesForCollections, default 50) is lifted so that the spec's per-node ratio decides. ItscodePointLimit(3 MiB) sits above this port's 2 000 000-character cap and never fires first. Its nesting cap is 1000 levels, as before; an alias chain can nest the materialized tree deeper than anything that can be written, and one deeper than 1000 levels is refused asdocument.limit.depthat$instead of being constructed. This guard is a port-specific addition beyond the spec, there to avoid a stack overflow in SnakeYAML's constructor when the alias maximum is raised. Up to 1000 levels the document model's own depth limit (200) decides, with the real path of the offending node; only above 1000 is the path$. - Complex keys. A mapping or sequence used as a key (
? [a, b]) is counted as if it were a value, so it cannot hide an expansion; the document model refuses such a key afterwards in any case. - Surfaces. Every YAML read goes through
YamlCodec:YamlCodec.read(reference defaults),YamlCodec.readWithLimits(configured), the CLI (defaults; there is no flag, and an over-limit input exits 2 withError: ...or, with--json, the code and path) and the conformance runner. The OML and OSD readers have no alias mechanism.
Programmatic schema rules (S-8, S-22, S-23, S-24, OSD-16)¶
Implemented in 0.4.0-alpha against omnist-spec v0.28.0-beta (commit 1a7d0de). No conformance vector
reaches any of them (DIV-5: OSD text arrives as declaration order and no text carries max = 0), so the
only pin is src/test/java/dev/omnist/schema/SchemaV028Test.java. A schema built through the Java API
(new Schema, new Record, new Type.Ref) is checked at construction, and the violation is a
dev.omnist.schema.SchemaException (an IllegalArgumentException) carrying getCode() and getPath():
| Rule | Code | Path | Raised by |
|---|---|---|---|
S-8: a record name, the root name or a Ref target is not [A-Za-z_][A-Za-z0-9_]* |
schema.invalid-name |
$, the name in the message only |
Record, Type.Ref, Schema constructors |
| S-22: a field label does not encode to valid UTF-8 (a lone UTF-16 surrogate) | schema.invalid-label |
the record name R; the label is never put in the path |
Record constructor |
OSD-16 / S-24: a field has max = 0 |
write.unsupported-value |
the record name R |
OsdWriter.write and OsdWriter.writeCompact (a WriteException) |
- S-23 (
schema.unknown-record) cannot arise. The Java API has no caller-supplied record ordering:Schema(root, records)takes one map and keeps its iteration order, so no ordering entry can name a record absent from it. [0,0]stays representable.new Field(label, type, 0, 0)is accepted (S-15, omnist-spec#83 is open); only the OSD writer refuses it.SchemaAlgebra.pruneremoves such fields (an unsatisfiable root record is kept intact, as the spec says),extractandnormalizenever produce one, soprunebefore writing is the way to clear it. There is no OSD-OML schema writer in this port (omnist-j#105), so the OSD-OML half of OSD-16 does not apply yet.- OSD text from a
String.OsdReader.readof text holding a lone surrogate reportsschema.invalid-labelat the record name as anOsdParseException. Input bytes are a different matter: a malformed byte sequence is decoded or refused before the reader sees it (D-14,parse.invalid-encoding). - Several violations. Construction reports the first one found; which is implementation-defined.
inferrecord names are ASCII. Because a record name is a Name (S-8),infernow maps every character outside[A-Za-z0-9_]to_when it derives a record name from a label (a label with an accented letter used to yield a non-ASCII name, which S-8 rejects), and an all-digit label yieldsRec. A malformedrootNamepassed toinferisschema.invalid-nameat$.
Known gaps¶
- D-14 (input must be valid UTF-8). The library API takes
String, so decoding is the caller's.Clidecodes both files and stdin strictly (CharsetDecoderwithREPORT) and refuses malformed input withparse.invalid-encodingat1:1; thebytes_hexvectors (E-27) cover it. - Codec syntax positions (JSON/YAML/TOML/XML) keep each codec library's own position arithmetic. E-28's code-point column is implemented for OML and OSD text only; the codec positions are the open question omnist-spec#114.
- Nesting past the parsers' own limits. JSON and YAML nesting of 1000 or more levels is refused
by Jackson / SnakeYAML before this port's depth limit (200) is consulted, and is reported as
parse.codec-syntaxrather thandocument.limit.depth. - Input-size guard. Input over 2,000,000 characters is refused with
document.parse-error(OML/OSD:parse.input-too-large/schema.input-too-large), codes that are not in the §8.3 taxonomy; the spec has no code for this cap. - OSD field label with a control character has no OSD spelling (OSD-14):
OsdWriterfails withwrite.unsupported-valueat the record path, unconditionally.
Testing¶
836 tests passing, 0 failures — JUnit unit/integration tests plus
jqwik property-based and fuzz tests (grammar-aware generators for TOML
radix literals, OML lexing, and YAML timestamp shapes; raw-input fuzzers
for every codec reader) run at thousands of iterations per property with
zero crashes. NoRawByteOrderMarkTest fails the build if any source file, tests included,
contains a raw U+FEFF.
Code coverage (JaCoCo)¶
Gate-scoped (excludes dev.omnist.conformance, the harness itself, and
CliMain, which is a thin argument-parsing entry point). Numbers are from
target/site/jacoco/jacoco.xml after a fresh mvn clean test, measured three times
(spec v0.28.0-beta adoption, 2026-10-02): the three runs gave identical figures.
| Package | Line | Branch |
|---|---|---|
| Overall | 99.71% (10 of 3443 missed) | 99.29% (17 of 2408 missed) |
dev.omnist.document |
100.0% | 98.6% |
dev.omnist.schema |
100.0% | 99.6% |
dev.omnist.algebra |
100.0% | 99.7% |
dev.omnist.cli |
99.4% | 98.0% |
dev.omnist.codec |
99.7% | 99.3% |
dev.omnist.validation |
100.0% | 100.0% |
dev.omnist.oml |
99.3% | 99.3% |
The CI gate (pom.xml) is set at 99.6% line / 99.1% branch.
v0.28.0-beta (2026-10-02): 836 tests; 10 of 3451 lines and 17 of 2414 branches missed, three runs identical
(headroom 3 lines and 4 branches). The new SchemaRules, SchemaException and writer paths are fully covered.
The margin is thin, and this is a known risk. At these numbers the gate has headroom of
only 2 more missed lines (13 allowed, 11 missed) and 3 more missed branches (21 allowed, 18
missed). The YAML alias limits added lines and branches (3214 to 3404 lines, 2228 to 2354
branches), and every new line and branch of YamlAliasCheck, YamlLimits and the new YamlCodec path is
covered: the missed counts are unchanged at 11 / 18. (A first draft of YamlAliasCheck left 3 lines and 7
branches uncovered, all defensive null or default arms that no input can reach; they were removed rather than
excluded, which brought the branch ratio from 98.94%, under the gate, to 99.24%.) The headroom was thinner at v0.21.0-beta than at
v0.19.0-beta (3 lines / 3 branches), because that sweep's new code (Cli.java's D-14 strict-UTF-8 decode and its --json structured-error
branches, Track2Runner's bytes_hex routing and its canonical-schema byte-for-byte comparison)
added lines faster than the new CliTest/Osd14Osd15/property-test coverage could close every
branch. Two consecutive mvn clean test runs gave identical numbers (11 missed lines, 18
missed branches both times), so this is not run-to-run jqwik seed noise -- it is a real,
reproducible margin that the next behavior-changing PR should watch closely. dev.omnist.cli
in particular dropped from 100.0%/98.9% to 99.4%/98.0% branch: one defensive catch
(Cli.java's MAPPER.writeValueAsString failure path inside the new --json error handler)
is confirmed-unreachable in practice (the JsonResponse/JsonError shapes serialize
unconditionally), the same class of documented trip-wire as the pre-existing gaps below.
The handful of remaining uncovered lines are documented trip-wires:
branches that are defensively correct but not reachable given the real
runtime behavior of the underlying libraries under this codebase's exact
configuration (verified empirically, not assumed) — e.g. a TOML parser
error path that can't be triggered without a malformed upstream parser
result, or an XML DOM check for a node type that setCoalescing(true)
already rules out before the DOM is exposed. Each is annotated in place
with the reasoning and, where relevant, the diagnostic that confirmed it.