PolicyPilot - Medical Benefit Policy Intelligence Portal
I built a citation-backed RAG portal that unifies fragmented multi-payer insurance policies so clinicians can answer drug coverage and prior-authorization questions without digging through 50-500 page PDFs.
01
TL;DR
- I built a citation-backed RAG portal that unifies fragmented multi-payer insurance policies so clinicians can answer drug coverage and prior-authorization questions without digging through 50-500 page PDFs.
- Best published result: 17 MCP tool APIs with Zod validation at every data boundary
02
Problem
- Clinicians manually dig through 50-500 page insurance payer policy PDFs to answer drug coverage/prior-authorization eligibility questions Clinicians and care teams handling prior-authorization and coverage eligibility decisions. Fragmented, unversioned payer policy data makes it easy to silently surface criteria from an outdated policy in a high-stakes healthcare decision.
03
My Role
- hackathon (Prompt Opinion hackathon) with defined system ownership
- Problem framing: Clinicians manually dig through 50-500 page insurance payer policy PDFs to answer drug coverage/prior-authorization eligibility questions
- Architecture: Express 5 REST API + 17-tool MCP server on a single Node.js process
- Implementation: PDF-to-structured-data pipeline (parse -> clean -> chunk -> validate) with Zod schema validation at data boundaries
- Evaluation: 17 MCP tool APIs with Zod validation at every data boundary; WCAG AA accessibility coverage
- Before: Fragmented, unversioned payer policy data makes it easy to silently surface criteria from an outdated policy in a high-stakes healthcare decision
- Personally designed: Architected an evidence-grounded RAG pipeline forcing citation-backed answers because Eliminates hallucination risk in high-stakes healthcare eligibility decisions; Designed a 17-tool MCP server (e.g. get_drug_coverage, check_patient_readiness, diff_policy_versions) exposing structured LLM tool-calling for PA extraction, patient-readiness evaluation, and cross-payer comparison because Lets external AI agents replace manual multi-payer PDF lookups with direct queries; Migrated from Ollama to Gemini 2.5-flash after benchmarking because Lower hallucination rates on policy-grounded questions; Built citation/version tracing into every LLM answer because Surfaces retrieval mismatches (e.g. an outdated policy version) immediately instead of silently misleading a provider
- Others owned: External datasets, APIs, academic baselines, or hackathon constraints shaped the work; the project page calls out what the source data verifies.
04
Constraints
- Built May 2026 (hackathon (Prompt Opinion hackathon)). Hackathon timeframe; healthcare data required citation-backed accuracy over speed.
05
Architecture
- Input: Multi-payer policy PDFs and patient documents
- Backend: Express 5 REST API + 17-tool MCP server on a single Node.js process
- Data & storage: PDF-to-structured-data pipeline (parse -> clean -> chunk -> validate) with Zod schema validation at data boundaries
- External APIs: Google Gemini 2.5-flash, FHIR (Da Vinci guides)
- Output: React SPA with unified drug coverage search across payers, cross-payer side-by-side comparison, real-time policy-change tracking, AI chat for coverage/PA/eligibility questions, and an Evidence Explorer for verifying answers against source documents
Express 5 REST API + 17-tool MCP server on a single Node.js process; PDF-to-structured-data pipeline (parse -> clean -> chunk -> validate) with Zod schema validation at data boundaries; React SPA with unified drug coverage search across payers, cross-payer side-by-side comparison, real-time policy-change tracking, AI chat for coverage/PA/eligibility questions, and an Evidence Explorer for verifying answers against source documents
- input 01Input
Multi-payer policy PDFs and patient documents
- process 02Backendinput ->
Express 5 REST API + 17-tool MCP server on a single Node.js process
- storage 03Data / storagebackend ->
PDF-to-structured-data pipeline (parse -> clean -> chunk -> validate) with Zod schema validation at data boundaries
- external 04External APIsbackend ->
Google Gemini 2.5-flash, FHIR (Da Vinci guides)
- output 05Outputstorage ->external ->
React SPA with unified drug coverage search across payers, cross-payer side-by-side comparison, real-time policy-change tracking, AI chat for coverage/PA/eligibility questions, and an Evidence Explorer for verifying answers against source documents
Routes
- Input -> Backend
- Backend -> Data / storage
- Backend -> External APIs
- Data / storage -> Output
- External APIs -> Output
06
Key Technical Decisions
- Architected an evidence-grounded RAG pipeline forcing citation-backed answers
- Designed a 17-tool MCP server (e.g. get_drug_coverage, check_patient_readiness, diff_policy_versions) exposing structured LLM tool-calling for PA extraction, patient-readiness evaluation, and cross-payer comparison
- Migrated from Ollama to Gemini 2.5-flash after benchmarking
- Built citation/version tracing into every LLM answer
07
Implementation
- Input layer: Multi-payer policy PDFs and patient documents
- Core system: Express 5 REST API + 17-tool MCP server on a single Node.js process
- Data layer: PDF-to-structured-data pipeline (parse -> clean -> chunk -> validate) with Zod schema validation at data boundaries
- External boundary: Google Gemini 2.5-flash, FHIR (Da Vinci guides)
- User output: React SPA with unified drug coverage search across payers, cross-payer side-by-side comparison, real-time policy-change tracking, AI chat for coverage/PA/eligibility questions, and an Evidence Explorer for verifying answers against source documents
08
What Broke / What Didn't Work
- Rejected: Unconstrained LLM chat responses without mandatory citations. Chosen path: Architected an evidence-grounded RAG pipeline forcing citation-backed answers.
- Rejected: A single monolithic chat endpoint without discrete tool boundaries. Chosen path: Designed a 17-tool MCP server (e.g. get_drug_coverage, check_patient_readiness, diff_policy_versions) exposing structured LLM tool-calling for PA extraction, patient-readiness evaluation, and cross-payer comparison.
- Rejected: Staying on Ollama for local/self-hosted inference. Chosen path: Migrated from Ollama to Gemini 2.5-flash after benchmarking.
- Rejected: Returning answers without source/version attribution. Chosen path: Built citation/version tracing into every LLM answer.
- Citation-backed grounding adds retrieval and validation overhead but is required for healthcare-safe answers
- Retrieval and document structure mattered more than the model choice - policy PDFs vary widely in layout, and the same rule can appear under different sections, headings, or tables
- Standards like HL7's Da Vinci FHIR guides only cover transmitting coverage details - keeping underlying multi-payer data correct and versioned was the harder problem
09
Results
- 17 MCP tool APIs with Zod validation at every data boundary - Eliminated a class of runtime type-mismatch bugs across all external tool calls - Eliminated a class of runtime type-mismatch bugs across all external tool calls - master-resume
- WCAG AA accessibility coverage - axe-core accessibility test coverage on the PDF-to-structured-data pipeline - axe-core accessibility test coverage on the PDF-to-structured-data pipeline - master-resume
10
What I'd Change Now
- Expand payer coverage beyond the initial dataset
- Add real-time policy-change alerting
- Broaden FHIR standard coverage
11
Stack
- React
- TypeScript
- Node.js
- Express
- Google Gemini
- MCP
- Docker
- Zod
- FHIR
- Railway
- Vitest
- axe-core (WCAG AA)
12
Links
- Source docs: 2-projects.json
Ask me about the trade-offs.
- Why this architecture boundary exists: Express 5 REST API + 17-tool MCP server on a single Node.js process
- How I evaluated Eliminated a class of runtime type-mismatch bugs across all external tool calls
- The hardest tradeoff: Citation-backed grounding adds retrieval and validation overhead but is required for healthcare-safe answers
- What I would change next: Expand payer coverage beyond the initial dataset