Syllabus PDFs to a Live Study Tracker: AI Extraction, a Data-Loss Bug, and Three Production Failures

A full-stack AI course tracker built for a college student in one sitting, including a silent data-loss bug caught by testing against real documents and a live production deploy debugged end to end

A college student heading into a new semester had the ordinary version of a real problem: five classes, five syllabi, each a different PDF or Word document with its own layout, one of them running an entirely wrong semester's schedule table. Nothing tracked what was actually due across all five at once, and nothing would until someone read every document by hand and typed it all in somewhere. The ask was correspondingly plain: upload the real file, get an accurate task board back, no manual re-entry.

Built end to end in one sitting: an Express API over a real Supabase Postgres schema, a two-pass Claude Haiku extraction pipeline (course metadata and grading structure in one pass, every dated or undated schedule item in the other), and a React board that groups the result by day with per-course filtering. Direct-to-storage uploads keep large PDFs off the API's own request body, and a second Claude pass reasons over the student's own notes to produce a running per-course summary and ranked exam-topic predictions.

The existing test suite ran entirely against synthetic fixtures, so before calling extraction done, all five of the real syllabi were run through the live pipeline instead of trusted on faith. Three of five came back with zero items extracted, not partial coverage; the whole upload silently quarantined every time. The shape of the failure was identical in all three: the code validated the model's entire response as one atomic batch, so a single malformed field anywhere in a 30-to-60-item response threw out every other correctly-extracted item sitting right next to it.

Each of the three failures had a genuinely distinct root cause once traced individually, not one bug wearing three disguises. One document had a real item correctly marked as date-precision "tba" that also carried a leftover due-date the model shouldn't have attached to it. One grading component genuinely never states a percent or point value anywhere in that syllabus, and the validation had no way to accept a component that simply is unweighted. The third document's response was truncating mid-batch against an 8192-token output cap that had never actually been checked against what the model could really produce, confirmed directly against the live Anthropic Models API, which reported a real ceiling eight times higher.

The fix replaced whole-batch validation with per-item validation: each item and grading component is checked independently, a bad one is dropped and logged rather than taking the rest of the batch down with it, and the existing one-shot retry is now reserved for a response that is actually structurally broken, not one messy item inside an otherwise-good response. Trust in the fix came from re-running the exact same three real documents through the live pipeline afterward, not from reasoning about the diff: 60, 31, and 49 items came back where zero had before, and the full 154-test suite (107 API, 47 frontend) stayed green throughout.

The AI summary feature carries its own honesty check, enforced in code rather than left to the prompt: every ranked exam-topic prediction has to cite a verbatim substring of something the student actually wrote in a note, or of the syllabus's own source text, before that topic is allowed to reach a validated response at all. A confident-sounding prediction with no real citation behind it is rejected the same way a malformed field is, not shown to the student as if it were grounded.

Shipping it surfaced three separate real production failures, each confirmed against the actual live URL rather than assumed fixed once the build went green. Vercel's Root Directory setting had silently rescoped the entire monorepo build to the API workspace alone, so the frontend never built at all. Fixing that surfaced a second failure that looked at first like a Node version mismatch, and bumping the pinned engine version and the dashboard's own Node.js Version setting to the newest option changed nothing, until the actual runtime stack trace showed a custom Rust-based Function runtime, not stock Node, that had simply never implemented the specific module-loading feature a shared file depended on, regardless of which Node "version" the dashboard claimed.

The real fix was converting that shared file to plain CommonJS instead of continuing to chase a platform version number, which immediately broke the frontend's own production build in a third, smaller way: the bundler only applies CommonJS compatibility to installed packages by default, not to a project's own local files. That one was a single bundler config line, not a rewrite. Each of the three fixes was confirmed the same way the extraction bug was, a real request against the live deployment, cookie and all, not a passing build log.

ReactViteExpressSupabasePostgreSQLClaude APIZodVercel