29 Sept 2026

Windows 11, and an honest scorecard

Ten measurements, reported as they came out

LLMKIT just picked up a sixth environment: Windows 11, alongside the five stacks it already spans.

The audit

A platform-parity audit compared how the kit behaved on macOS with how it behaved on Windows. It surfaced seven gaps, across:

  • CLI scripts
  • Scheduling
  • Compilation
  • Update detection

We fixed what the audit found — and then verified against Windows 11 directly, instead of assuming the fixes held.

The results

Ten measurements:

  • 6 passed cleanly.
  • 3 came back “not measured” — Copilot, Codex and uninstall cleanup weren’t fully testable in that pass.
  • 1 was partial — 4 of 6 feeds pre-scheduled as expected.

Why report it this way

We could have rounded that up to “Windows support: done.” We didn’t.

“Not measured” isn’t the same as “passed,” and “partial” isn’t the same as “working.” Collapsing them hides exactly the information the team needs to decide what to test next. The honest version — what passed, what didn’t get measured, what’s still partial — is more useful than a clean bill of health that isn’t quite true yet.