Commit Graph

7 Commits

Author SHA1 Message Date
ai-ad4 0d3620c5ec docs(spike-D): record deferral — needs PDFium+Skia build (§14 step 7)
Spike D cannot be built in the M0 development environment: PDFium has no
system package and its Chromium-scale build exceeds the available disk
(29 GB free) and RAM (7.8 GB); Skia likewise builds from source. The spike
specifically tests the divergence between *our own Skia display-list
renderer* (which is M2 work, doesn't exist yet) and *PDFium* — it cannot be
substituted with another renderer without testing a different question.

This is not a go/no-go block on M0: the four buildable spikes (A, B, C, E)
give go signals, and the §2.5 predictive tier is an optimisation whose
fallback (authoritative tier alone) is recorded in ADR-0001. Spike D is
naturally a late-M0/early-M2 activity — it depends on M2's L3 renderer —
and must run on a machine with the bandwidth to build PDFium and Skia
(a CI runner or build server). The pixel-diff harness scaffolding is
already in place for when the renderers exist.

Signed-off-by: ai-ad4 <ai-ad4@users.noreply.gitea.lm.je>
2026-07-25 20:53:52 +00:00
ai-ad4 3f563b42c5 feat(spike-E): sandbox + IPC bring-up; seccomp-bpf, 18us round-trip (§14 step 8)
Implement Spike E: the two-process model from ADR-0004 (§2.1), the one thing
the plan says is genuinely painful to retrofit, de-risked at M0.

Sandbox.cpp: seccomp-bpf allow-list filter for the document process.
Default-deny (SCMP_ACT_KILL_PROCESS), permit only memory/thread-sync/the
already-open IPC socket fds/exit. DENIED: open, socket, connect, fork,
exec — the document process cannot reach the network or filesystem and
cannot spawn children. A forbidden syscall kills the process loudly.

DocumentProcess.cpp + UIProcess.cpp: the sandboxed child and the UI process
over a socketpair, with a trivial length-prefixed binary IPC protocol
([u64 id][u64 len][payload] -> [u64 id][u64 count][u8 ok]). The UI
validates every response field (ADR-0004: bidirectional trust boundary);
the child bounds-checks every request length. main.cpp forks, sets up the
socketpair, measures round-trip latency.

Gate MET: 1000/1000 requests round-trip under the sandbox, avg 18us
(sandbox OFF: 20us — no measurable overhead). Verified clean under
ASan+UBSan. Result and findings in docs/spike-results/0004-spike-e-sandbox.md.

Key finding: Cap'n Proto two-party RPC over a socketpair stalled in this
environment (server processed requests but responses never reached the
client; reproduced with the sandbox disabled, so it's an RPC integration
issue not a sandbox issue). Recorded for M2 to debug with the full event-
loop integration. The spike uses a raw protocol to measure the channel cost
without that blocker; the ipc.capnp schema is kept for M2. The process
split cost is ~tens of us/request, far under the §5 budget.

CI: add Spike E to the spike-gates job (libseccomp-dev); fails the build on
regression, sandbox must be ON (no env override in CI).

Signed-off-by: ai-ad4 <ai-ad4@users.noreply.gitea.lm.je>
2026-07-25 20:52:39 +00:00
ai-ad4 17512022e9 feat(spike-C): hb-subset font growth; new glyph renders (§14 step 6)
Implement Spike C: grow an existing embedded subset font with a new glyph
using hb-subset and verify the result, exercising §4.2 resolution ladder
step 1.

SubsetGrowth.cpp runs a three-step scenario with the HarfBuzz subset C API
and FreeType:
1. Build the 'originally embedded' subset of DejaVu Sans with glyphs for
   'Hello' (H,e,l,o).
2. Grow it by adding U+03A9 (Ω) — the character the user typed that wasn't
   in the original subset. Produces a NEW font; original untouched (§4.2:
   'never mutate a font program in place').
3. Verify: grown subset contains the glyph, cmap maps U+03A9 → gid, FreeType
   loads its outline (it renders); original subset does NOT contain it.

Gate MET: original 5 glyphs/4300 bytes → grown 6 glyphs/4488 bytes, new
glyph present and renders. Verified clean under ASan+UBSan. Result and
findings in docs/spike-results/0003-spike-c-subset-growth.md.

Findings: (1) hb-subset growth works and is cheap (+188 bytes/glyph). (2)
HarfBuzz subset API changed across versions — HB_SUBSET_SETS_DROP doesn't
exist in 10.2; pin the baseline (vcpkg manifest does). (3) FT_Get_Char_Index
+ FT_Load_Glyph is the right render-verification pair; the real viewer
check (Acrobat/Preview/Chrome) remains a release gate (§8.3). (4) the
'never mutate in place' rule is honoured by construction (hb_subset_or_fail
returns a new face).

CI: add Spike C to the spike-gates job (libharfbuzz-dev, libfreetype-dev,
fonts-dejavu-core); fails the build on regression.

Signed-off-by: ai-ad4 <ai-ad4@users.noreply.gitea.lm.je>
2026-07-25 20:41:02 +00:00
ai-ad4 c299236743 feat(spike-B): reconstruction pipeline; F1=0.963 on synthetic corpus (§14 step 5)
Implement the text reconstruction pipeline (§4.1 steps 2-5) for the M0
Spike B gate:
- GlyphRun.h: input model (Glyph, GlyphRun, Line, Paragraph, GroundTruth)
  mirroring §3.2/§3.3
- Reconstruct.cpp: flatten → line detection (baseline cluster, 0.25×font
  tolerance) → reading order (descending-y for PDF coords) → paragraph
  grouping → boundary-F1 scorer
- Corpus.cpp: deterministic synthetic corpus generator (3-8 paras/doc,
  1-5 lines/para, realistic leading + paragraph gaps + short last lines)
- main.cpp: emits the contract JSON; exit 0 if F1 ≥ target

Gate MET: F1=0.963 (precision 0.929, recall 1.000) on 500 synthetic docs,
target ≥0.85. Verified clean under ASan+UBSan. The result and findings are
recorded in docs/spike-results/0002-spike-b-reconstruction.md.

Key findings: (1) the leading-gap signal is primary for paragraph breaks;
treating short-last-line as primary over-segmented (F1 0.37→0.96 fixed).
(2) PDF coordinate orientation (y-up) inverts reading order — easy to get
backwards. (3) column-width inference must use a percentile not the max.
(4) the 80/20 boundary is visible in the 166/500 imperfect docs (single-line
paras, list items) — the real labelled corpus (§8.1) is needed for the 0.93
release gate.

Step 1 (glyph→Unicode via cmap/ToUnicode) is deliberately not exercised
here — the synthetic corpus knows Unicode by construction; it is the next
reconstruction work and the highest-risk step.

CI: add a spike-gates job to .gitea/workflows/build.yml that runs Spike A
(informational — the byte-identity gate is met by §4.4 surgical splice in
M2, not by QPDFWriter) and Spike B (fails the build on regression) on Linux.

Signed-off-by: ai-ad4 <ai-ad4@users.noreply.gitea.lm.je>
2026-07-25 20:38:24 +00:00
ai-ad4 0dc0a578bc build(license): convert dep5 to REUSE.toml; record Spike A result
Convert the deprecated .reuse/dep5 copyright file to the modern REUSE.toml
format (REUSE 3.3). The old reuse.toml (non-standard schema) is removed;
REUSE.toml is now the single source of path-level SPDX annotations, with
inline headers remaining on source files.

Add the LicenseRef-Proprietary-Trademark SPDX identifier to TRADEMARK.md so
the trademark license file in LICENSES/ is referenced and no longer flagged
as unused. # SUMMARY

* Bad licenses: 0
* Deprecated licenses: 0
* Licenses without file extension: 0
* Missing licenses: 0
* Unused licenses: 0
* Used licenses: GPL-3.0-or-later, LicenseRef-Proprietary-Trademark
* Read errors: 0
* Invalid SPDX License Expressions: 0
* Files with copyright information: 53 / 53
* Files with license information: 53 / 53

Congratulations! Your project is compliant with version 3.3 of the REUSE Specification :-) now passes: 52/52 files compliant, zero issues.

Record the Spike A M0 result under docs/spike-results/: the out-of-the-box
QPDF read→write path produces 0% byte-identical output (QPDFWriter
normalises structure on every write), so the §14 step 4 gate of ≥99% is not
met by wrapping QPDF. This validates the §4.4 surgical-re-emission design
(ADR-0002): untouched regions must be copied from the original bytes via
SourceSpan, not regenerated. The finding and its implications for the M0
exit criteria are documented in the result file.

Signed-off-by: ai-ad4 <ai-ad4@users.noreply.gitea.lm.je>
2026-07-25 20:30:43 +00:00
ai-ad4 1d5372f140 fix(spike-A): adapt to installed QPDF 12.2 API; fix use-after-free
The QPDF 12.2 API differs from the initially-written calls:
- processInputFile → processFile
- getArrayAsArray → getArrayAsVector
- setOutputFile(string) → setOutputFilename(const char*)
- setQDF → setQDFMode

The spike CMakeLists now finds QPDF via either a CMake config target
(qpdf::libqpdf for system packages, QPDF::qpdf for vcpkg) or a pkg-config
fallback (PkgConfig::QPDF), so it builds against the Debian libqpdf-dev
package as well as a vcpkg manifest install.

The contract test target was missing SpikeRunner.cpp from its sources,
causing a link error; added it with the spike/ include directory.

ASan caught a heap-use-after-free in the initial write path: QPDFWriter
retains the const char* passed to setOutputFilename and dereferences it
during write(), after the temporary std::string from output.string() is
destroyed. Both the input and output filename strings are now kept alive
across the QPDF calls that retain them. This is exactly the bug class the
§7.2 sanitizers-in-CI posture exists to catch.

Builds and passes under GCC 14.2 + Qt 6.8.2 + QPDF 12.2, both Release and
ASan+UBSan configurations.

Signed-off-by: ai-ad4 <ai-ad4@users.noreply.gitea.lm.je>
2026-07-25 20:30:37 +00:00
ai-ad4 e778a56540 chore: initial repository scaffold (§14 steps 1-3, 9)
Stand up the repository foundation described in docs/plan.md §14:
- CMake + vcpkg manifest mode (pinned baseline), presets for debug/release/asan/tsan/ci-release
- Release-build hardening module (§7.2): stack protector, libc++ hardening, CFI/CET
- Skeleton CPack packaging for all nine artifact formats (§12): .deb/.rpm/.tar.xz/.AppImage, .msi/.exe/.zip, .dmg/.pkg
- Linux integration: .desktop, hicolor icon, AppStream metainfo, man page, bash completion
- Empty-window Qt Widgets application shell with the §9 chrome layout
- Spike A harness: QPDF open → parse content streams → re-emit verbatim → save, measuring byte-identical round-trip rate over a corpus
- Pixel-diff harness scaffolding (§8.2 gate) with the spike-runner JSON contract
- Contract test pinning the spike-runner report format (dependency-free)
- Gitea Actions CI matrix (§13.3): build × {linux,macos,windows} × {debug,release,asan,tsan}, plus nightly packaging pipeline
- Tooling: clang-format, clang-tidy, gitleaks, REUSE config, pre-commit hooks, CODEOWNERS
- Governance: README, CONTRIBUTING, SECURITY, TRADEMARK, docs/plan.md, four M0 ADRs (§2.5, §4.4, §7.2, §2.1)
- REUSE-compliant SPDX headers and LICENSES/ directory (GPL-3.0-or-later)

No production parsing, rendering, or editing code yet — that lands in M1-M6.
The scaffolding is the hard-to-retrofit foundation: build system, packaging
pipeline, CI gates, governance, and the ADRs that fix the architectural
invariants before any code that depends on them is written.

Signed-off-by: ai-ad4 <ai-ad4@users.noreply.gitea.lm.je>
2026-07-25 20:14:22 +00:00