Class Bom

java.lang.Object
dev.omnist.document.Bom

public final class Bom extends Object
The one place the leading byte-order mark rules live (omnist-spec section 2.5, D-15 and D-21).

Every read surface (OML, OSD, JSON, YAML, TOML, XML) calls strip(java.lang.String, java.util.function.Supplier<? extends java.lang.RuntimeException>) on its input text before doing anything else, so the rule is stated once rather than re-implemented per reader:

  • D-15. A leading U+FEFF (offset zero) is consumed and contributes nothing.
  • D-21. Exactly one is consumed. A second U+FEFF still standing at offset zero of the text that remains is rejected, reported at text position 1:1 of the remaining text. No library gets a chance to swallow it quietly.
  • A U+FEFF anywhere else is ordinary content and is never touched.

Writers never emit a mark (D-15), so there is no writer-side counterpart here.

The mark is spelled (char) 0xFEFF in code and as a Java unicode escape in tests, never as a raw invisible character in a source file; NoRawByteOrderMarkTest enforces that.

  • Field Details

  • Method Details

    • strip

      public static String strip(String text, Supplier<? extends RuntimeException> onDoubled)
      Strips exactly one leading byte-order mark and rejects a second (D-15, D-21).
      Parameters:
      text - the input text; null is returned unchanged
      onDoubled - supplies the surface-specific exception to throw when a second U+FEFF is still at offset zero after the strip
      Returns:
      text without its single leading mark, or text itself if it had none