What is the user buying?
Not chat. A trustworthy resolution.
When a support question touches payments, APIs, policies or account behavior, a fluent answer is not enough. Proofline retrieves approved evidence, shows what supports the answer, refuses when the evidence is insufficient, and hands a complete case to a human when the question is account-specific or high-risk.
This case goes deeper into knowledge lifecycle, retrieval quality, refusals and how trustworthy support operates six months after launch.
When a support question touches payments, APIs, policies or account behavior, a fluent answer is not enough. Proofline retrieves approved evidence, shows what supports the answer, refuses when the evidence is insufficient, and hands a complete case to a human when the question is account-specific or high-risk.
The five-step reasoning pattern stays consistent; the actual product logic is specific to this problem.
Not chat. A trustworthy resolution.
The answer must be applicable, current, cited and safe to act on.
Bad source, weak retrieval, stale version, PII, prompt injection, overconfidence.
Interpret, retrieve, synthesize and explain; policy controls scope and escalation.
Failures become corpus fixes, retrieval changes, prompt changes or new regression cases.
This is how I would reuse the thinking without copying the product.
Same evidence/grounding pattern; clinical safety and regulation tighten the autonomy boundary.
Same citation and versioning problem; jurisdiction and precedent become critical metadata.
Same retrieval problem; access control and confidential data dominate.
Exact API versions, identifiers and code compatibility matter more than broad semantic similarity.
AI products often fail when one persona is treated as ‘the user.’
Wants a verifiable answer quickly, without needing to understand retrieval or model behavior.
Needs the same evidence, classification and conversation summary when escalation occurs.
Wants safe deflection, lower handling effort and visibility into knowledge gaps without increasing wrong-answer risk.
User states the question.
PII/scope gate runs first.
Find exact + semantic evidence.
Version/freshness filters remove invalid context.
Decide if the question is answerable.
Confidence is not only model probability; evidence quality matters.
Generate cited response or explain why the system cannot safely answer.
Account-specific questions prepare handoff.
Feedback + ticket outcome become knowledge/eval input.
Fix corpus, retrieval or policy—not blindly the prompt.
Each decision includes an implicit reversal test: better evidence can change the choice.
If the evidence cannot support the claim, the product should preserve trust rather than maximize deflection.
Error codes and API identifiers need exact retrieval; natural-language questions benefit from semantic retrieval.
Source ownership, freshness, deprecation and conflicting documents determine quality months after launch.
Public documentation can support policy-level answers; account data should require authenticated tools and stronger controls.
Every connector has a defined job. Animated dashed lines represent active evaluation/learning loops rather than decorative motion.
Parse approved HTML/PDF/API schema; chunk at structure boundaries; attach URL, section, version and last-modified metadata.
Detect scope, PII, exact identifiers and whether the question needs public knowledge or account state.
Lexical lane for codes/IDs + dense semantic lane for meaning; merge and metadata-filter.
Cross-encoder reranks the short list so the model sees the most applicable evidence first.
Check source coverage, freshness, contradictions and scope before generation.
Produce short answer + steps + citations; only claims supported by context are allowed.
Faithfulness/citation checks + deterministic policy decide Answer / Clarify / Escalate / Refuse.
Human corrections and unresolved query clusters trigger doc audits, new evals and re-indexing.
Component eval, system eval and product outcome are deliberately separated.
Precision@K / recall@K, reranker relevance, citation resolution, PII redaction, source freshness.
Faithfulness, correct refusal, scope routing, handoff completeness and latency/cost.
Prompt injection, malicious retrieved text, unauthorized tool attempt and data leakage.
Ticket deflection without repeat contact, satisfaction, escalation quality and unresolved knowledge-gap rate.
Offline eval → internal support alpha → shadow comparison → limited beta → monitored GA.
Thumbs-down/ticket outcome → failure taxonomy → fix source/retrieval/prompt/policy → rerun suite.
The PRD is intentionally feature-level and includes AI behavior, deterministic controls, telemetry and non-goals.
A merchant asks why a settlement is delayed. Public policy can explain the general cycle, but the system must not invent account-specific status.
Answer policy-level settlement questions from approved sources and identify when authenticated account data or human support is required.
query_received · pii_redacted · retrieval_completed · citation_rendered · answer_refused · escalation_created · thumbs_down · ticket_within_24h
Typed state/schema · API/tool contracts · error states · permissions · eval fixtures · analytics events · rollout/rollback.
Use standard infrastructure for storage, models, tracing and connectors where it does not create strategic advantage.
Timeouts, stale data, partial results, rate limits and provider failures are explicit product states—not invisible backend details.
Every click represents a user/product decision and explains why the information is needed.
Proofline does not begin by generating. It begins by deciding what evidence the question requires.
Exact terms and policy headings.
Meaning-based retrieval.
Old policy versions removed.
Holiday exception ranks above generic cycle page.
This prototype only demonstrates policy-level grounding. It does not claim to know this merchant's actual settlement status.
Short answer first; uncertainty preserved.
Page + section + version metadata shown.
Public corpus cannot reveal private transaction state.
Sanitized transcript + evidence package.
A quick signal helps me understand what is useful to recruiters, founders and product teams.
Prove approved sources are complete enough.
Support team compares AI evidence to real resolutions.
Public-doc questions only; visible citations + escalation.
Read-only account tools with stronger permissions and evals.
Gap detection, doc-owner workflows and continuous quality management.
Pivot or stop if the approved corpus cannot be kept current enough to support safe answers, or if users still require the same human investigation after seeing the grounded response.
Where the source material does not prove an implementation, the portfolio says proposed/candidate rather than “built with.”
NIST frames trustworthiness and risk management across the AI lifecycle. Product implication: risk controls belong in design, evaluation and operations—not only at launch.
Open source ↗OWASP notes that RAG does not eliminate prompt-injection risk. Product implication: retrieved content and user prompts need explicit trust boundaries and authorization controls.
Open source ↗The source PRD defines 500-token/50-overlap chunking, text-embedding-3-small, Cohere rerank-v3, Claude Haiku/Sonnet roles, pgvector, 0.7 routing and cited/refusal behavior. Proofline generalizes the reasoning without implying an official product.
Internal/source-project basisNo spreadsheet preview is embedded. Open the workbook only if you want the detail.