Lydia Ng
05 · Case Study / zephyr-scripts/opinionbot-v2
AI Product Prototype & System

OpinionBot

Retrieval-Augmented Knowledge & Analytics System

RAG Knowledge Management Web Ingestion Analytics
“From scattered web content to a grounded conversational knowledge layer.”
Dual-Track System Architecture

Ingestion, Retrieval & Observability Pipeline

Primary Knowledge Path Observability Bus
Track 01 · Production Ingestion & RAG Orchestration Deterministic chunking · Pinecone top-k · Grounded synthesis
01 / INGRESS

Web Sources

Public URLs, essays, structural whitepapers, opinion articles.

link HTTP / HTTPS
02 / EXTRACT

Ingestion Engine

Puppeteer crawler, Axios fetcher, HTML hygiene & Markdown AST parser.

cleaning_services Sanitized MD
03 / TRANSFORM

Chunk & Embed

Deterministic sliding window (512 tk) with OpenAI text-embedding-3.

splitscreen 1536-dim Vec
04 / INDEX

Pinecone Index

Partitioned vector space indexed across strict domain namespaces.

database Namespace Isolation
05 / RETRIEVE

Hybrid Search

Cosine top-k vector similarity with semantic reranking filters.

filter_alt Top-K Context
06 / GROUND

Context Stamping

Provenance stitching & prompt assembly with explicit source headers.

verified Citations Match
07 / GENERATE

Grounded AI

Synthesized answer strictly bounded by retrieved passage citations.

psychology Streamed Output
ASYNC TELEMETRY & FEEDBACK PIPELINE
Track 02 · Observability, Sentiment & Knowledge Gap Analytics PostgreSQL log bus · Aggregated telemetry · Semantic heatmaps
DATA LAYER table_rows

Supabase SQL Store

Persists complete conversation logs, prompt variations, retrieved similarity metrics, and user verification votes (helpful / ungrounded).

pgvector-ready RLS policies
SEMANTIC COMPUTATION bubble_chart

Topic & Sentiment Extraction

Periodic batch clustering of user queries to group inquiries into topical clusters, tone profiles, and emerging intellectual inquiries.

K-Means VADER / LLM Tone
OBSERVABILITY CONSOLE monitoring

Real-time Analytics Console

Executive dashboard exposing query frequencies, low-confidence queries (gap discovery), domain coverage heatmaps, and latency metrics.

Gap Discovery P95 < 1.2s
Section 01

The Challenge

Unstructured knowledge & hallucination risks

Institutional thinking, critical analysis, and strategic intelligence are disproportionately locked inside lengthy opinion essays, external commentary, and sprawling documentation. Traditional search engines only match keywords without synthesizing arguments, forcing analysts into tedious manual cross-referencing.

The Keyword Failure

Keyword engines fragment context. When an analyst asks about multi-party geopolitical stances or subtle regulatory positions, boolean search returns hundreds of disjointed links rather than a structured synthesis.

The Generic LLM Liability

Off-the-shelf generative models speak with confident fluency but lack source provenance. In high-stakes environments, ungrounded responses constitute unacceptable epistemic risk.

Section 02

The Thinking

A 5-tier architectural separation of concerns

Solving the synthesis dilemma required resisting monolithic prompts. We designed an explicit 5-tier pipeline where content acquisition, retrievable storage, context filtration, generative synthesis, and usage intelligence function as isolated, testable modules.

01 Acquisition & Hygiene Extracts DOM bodies, removes navigation noise, boilerplate, and tracking scripts, normalizing prose into clean Markdown AST nodes.
02 Namespace Vector Storage Partitions text chunks with strict domain-scoped namespaces in Pinecone, preventing cross-contamination between publication sources.
03 Dynamic Query Retrieval Translates user intent into vector coordinates, applying threshold filters to avoid pulling low-affinity, misleading content into the prompt.
04 Strict Citation Synthesis Enforces deterministic system prompts that command the LLM to only synthesize statements that directly map to indexed passage stamps.
05 Aggregate Telemetry Treats conversational queries as market research data, extracting organizational intelligence gaps when queries yield low-confidence hits.
Section 03

The Solution

Functional modules & user interface capabilities

OpinionBot delivers a unified interface pairing real-time knowledge ingestion with citation-first conversational exploration, backed by an operator dashboard for knowledge base observability.

language

Automated Content Ingestion

Dual-strategy crawler engine: lightweight Axios scraping for static editorial layouts and headless Puppeteer execution for dynamic, JavaScript-hydrated content sites.

Features: URL queueing · DOM pruning · Metadata extraction
folder_copy

Namespace-Isolated Vectors

Partitioned vector indexing in Pinecone allows granular source scoping. Users can direct queries toward specific essayists, publishers, or aggregate domains simultaneously.

Features: Dynamic namespace routing · Metadata tags
quick_reference

Grounded Citations

Every generative assertion contains inline interactive badges linked directly to the crawled sentence fragment, source URL, and timestamped extraction passage.

Features: Exact fragment match · Zero hallucination constraints
history_edu

Query History & State

Session continuity and user authentication via Firebase Auth and Supabase. Analysts can bookmark synthesis threads, export executive digests, and compare citations.

Features: Auth tokens · Markdown thread export · Query tagging
analytics

Institutional Analytics Dashboard

Continuous Evaluation

Rather than treating the chat interface as an ephemeral terminal, OpinionBot feeds query telemetry into an analytical dashboard. The system categorizes incoming questions, measures user sentiment, and flags queries where vector similarity falls below threshold—explicitly signaling to content teams where institutional literature is lacking.

METRIC 01Topic Frequency
METRIC 02Retrieval Confidence
METRIC 03Sentiment Polarity
METRIC 04Knowledge Gaps
Section 04

System & Tech Stack

Rigorous, multi-tier service orchestration

The infrastructure avoids speculative frameworks in favor of resilient production primitives: TypeScript contracts from end-to-end, high-performance vector lookup, and partitioned persistent data stores.

FRONTEND LAYER

Client Experience

  • • Vite & React (SPA)
  • • TypeScript Strict Mode
  • • Tailwind CSS Design System
  • • Real-time SSE Stream Consumer
SERVICE RUNTIME

Backend Execution

  • • Node.js & Express Engine
  • • LangChain Pipeline Controls
  • • Dynamic Rate Limiting
  • • Async Worker Job Queues
AI & VECTOR

Semantic Compute

  • • OpenAI text-embedding-3
  • • GPT-4o Guided Generation
  • • Pinecone Cloud Vector DB
  • • Namespace Routing Logic
CONTENT PIPELINE

Ingestion & Hygiene

  • • Puppeteer Headless Cluster
  • • Axios HTTP Client
  • • Cheerio HTML AST Normalizer
  • • Window Chunking Tokenizer
DATA PERSISTENCE

Application Database

  • • PostgreSQL via Supabase
  • • Row-Level Security Policies
  • • Analytics Telemetry Logs
  • • User Session Context Store
SECURITY & IDENTITY

Identity Guard

  • • Firebase Authentication
  • • JWT Token Verification
  • • Sandboxed Scraping Workers
  • • Ephemeral Context Sanitation
Section 05

Outcome & Demonstration

Validation of production principles

An extensible knowledge interface proving how unstructured web information transforms into an auditable intelligence layer.

OpinionBot validated that knowledge systems succeed not through larger foundation model parameters, but through deterministic ingestion hygiene, strictly scoped retrieval boundaries, and continuous observational feedback loops.

Architectural Principles Demonstrated
01

Production RAG Architecture with Provenance Guarantees

Elimination of synthetic fabrication by restricting model generation strictly to retrievable context chunks accompanied by deterministic citation metadata.

02

Web-Scale Content Ingestion & AST Hygiene

Engineering resilient scraping workers capable of handling unpredictable web formatting, removing DOM detritus, and maintaining content integrity before embedding.

03

Multi-Tier Service Orchestration

Seamless synchronization across asynchronous crawlers, managed vector databases (Pinecone), relational logging (Supabase PostgreSQL), and client interfaces.

04

Macro Analytics Over Conversational Interactions

Transforming individual user queries into aggregate institutional telemetry to systematically uncover intellectual blindspots and high-demand research topics.