Files
accounted/lib/reports/sru-encoding.ts
T
Jakob Wennberg ec27228a8e style: remove em/en dashes repo-wide, add CLAUDE.md rule against them (#890)
Em dashes (—) and en dashes (–) had spread across comments, docs, tests,
and a few UI strings, reading as AI-generated boilerplate rather than
house style. Replaced each with punctuation matching its context: colon
for explanatory clauses, comma for asides, plain hyphen for numeric/legal
ranges (e.g. "21-23§"), "to"/"till" for date ranges, parentheses for
paired-dash asides. messages/en.json and messages/sv.json were fixed by
hand together to keep sv/en in sync.

Left untouched where the dash is the functional subject rather than
decorative punctuation: date-range-parser.ts's separator regex,
charset-repair.ts's CP1252 byte-mapping table (and its test), the SIE
encoding mojibake docs, generic-csv.ts's minus-sign normalizer, the
agent system-prompt files that already instruct against em dashes, and
a golden iXBRL test fixture compared byte-for-byte.

Also fixes two bugs surfaced along the way: an off-by-one in
ApiKeysPanel's scope-label split (a leftover from an earlier partial
pass), and a charset-repair test that had lost the literal en-dash it
exists to verify.

Regenerated the agent atom seed migration (skills:generate) since 27
SKILL.md files changed. Added a CLAUDE.md rule against em/en dashes,
with an explicit carve-out for the functional-dash cases above.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-04 15:58:06 +02:00

22 lines
732 B
TypeScript

/**
* Shared encoding helpers for Skatteverket SRU files.
*
* SRU submissions (INFO.SRU + BLANKETTER.SRU) must be ISO 8859-1 (Latin-1),
* never UTF-8: Swedish characters (å, ä, ö) corrupt otherwise and Skatteverkets
* filöverföringstjänst rejects the upload. This is the single most common cause
* of programmatic SRU validation failure.
*/
/**
* Encode a string as ISO 8859-1 (Latin-1) bytes.
* Characters outside the Latin-1 range (> 0xFF) are replaced with '?' (0x3F).
*/
export function encodeISO88591(str: string): Uint8Array {
const bytes = new Uint8Array(str.length)
for (let i = 0; i < str.length; i++) {
const code = str.charCodeAt(i)
bytes[i] = code <= 0xff ? code : 0x3f
}
return bytes
}