Skip to content
Knowledge Assistant
Contact
AIAI assistantIn build

Knowledge Assistant

A RAG assistant that answers from company docs, with citations.

Client
Consulting group (NDA)
Role
AI engineer
Year
2026
Industry
Professional services
Built withNext.jsTypeScriptpgvectorLLMLangGraphRedis

Problem

The answer existed, somewhere

Policies, SOPs and project knowledge were spread across SharePoint, Google Drive and Confluence. People asked colleagues instead of searching, and the same questions came up again and again.

BeforeEmail, sheets and chat
  • Does anyone know where the latest travel policy is?Operations channel, via chat
  • Re: Which onboarding checklist do we use now?Consultant, via email
  • Expenses policy (old) - copy.docxGoogle Drive, via doc
  • Who approves client-billed travel?Operations channel, via chat
Representative examples, reconstructed from discovery interviews.
  • Scattered sources

    SharePoint, Drive and Confluence each held part of the answer, with different search and access rules.

  • Keyword search misses

    Searching for 'travel approval' didn't find a policy titled 'Expenses and bookings'.

  • Interrupt-driven answers

    Senior staff answered the same questions in chat, repeatedly, instead of doing their own work.

  • Stale copies

    Downloaded copies of old policies kept circulating long after the originals had changed.

Research

Finding where answers actually live

Before choosing a retrieval approach, we mapped where knowledge lives, who is allowed to see it, and which questions people ask most. Permissions turned out to be the hard part, not the model.

  • Question log

    Collected the questions people asked in team channels and help requests, grouped by topic.

  • Source inventory

    Listed SharePoint sites, Drive folders and Confluence spaces, with their owners and access groups.

  • Permission mapping

    Mapped how each source expresses access, so retrieval could respect it exactly.

  • Answer review

    Knowledge managers graded draft answers to agree what 'good enough to ship' means.

What we learned

  • Most questions already had an answer; people couldn't find it, or didn't trust it was current.
  • An answer without a source wasn't trusted, however fluent it sounded.
  • Permissions differ per source, so they must be synced with the content, never assumed.
  • Unanswered questions are valuable in themselves: they show where documentation is missing.

User journey

From 'who knows this?' to a cited answer

A composite persona from the consulting team, following one policy question.

BeforeWith the product

  1. 1

    Ask a question

    Before: Guessed keywords across several search boxes.

    Now: Asks in plain language, in one place.

  2. 2

    Find the right document

    Before: Opened promising files that turned out to be out of date.

    Now: Retrieval ranks passages from current documents he is allowed to see.

  3. 3

    Trust the answer

    Before: Messaged a senior colleague to double-check.

    Now: Every sentence cites the passage it came from, one click away.

  4. 4

    Go deeper

    Before: Started a new search for each related policy.

    Now: Follow-up questions keep the context of the conversation.

  5. 5

    Close the gap

    Before: Unanswered questions disappeared into chat history.

    Now: A thumbs-down or 'no answer' shows knowledge managers what to document next.

Solution

Answers with receipts

A retrieval-augmented assistant that searches only what the signed-in person is allowed to see, answers in plain language, and cites the exact passages it used. A LangGraph flow decides when to search, when to ask a clarifying question and when to say it doesn't know.

  • Permission-aware retrieval

    Access groups are synced with every document and applied inside the search query itself.

  • Inline citations

    Each claim links to the passage it came from, with a preview of the source on hover.

  • Knows when it doesn't know

    If nothing relevant is found, it says so instead of guessing, and logs the gap.

  • Feedback loop

    Thumbs up or down on every answer, reviewed by knowledge managers in an admin view.

How the work flows

Switch between the old process and the one the product runs.

  1. 1

    Ask in plain language

    One place, natural questions.

    Person, in productOwner: Consultant
  2. 2

    Retrieve from permitted sources

    Search is filtered by the person's access groups.

    AutomatedOwner: Answer graph
  3. 3

    Answer with citations

    Every claim links to its passage.

    AutomatedOwner: Answer graph
  4. 4

    Verify in one click

    Open the cited passage in its source.

    Person, in productOwner: Consultant
  5. 5

    Feedback flags gaps

    Unanswered questions reach knowledge managers.

    AutomatedOwner: Admin view
One question, one cited answer, and every gap is recorded for someone to fix.

Architecture

Permissions travel with the content

Connectors sync documents and their access groups together. Chunks are embedded and stored in pgvector alongside those groups, so the answer graph can filter by the signed-in person's access inside the query. Restricted text never reaches the model.

  • Interface
  • Service
  • AI
  • Data and queues
  • External system

Sources

Ingestion

Knowledge

Answering

Answer graph

LangGraphAI

Rewrites the question, retrieves with a permission filter, re-ranks, then answers or abstains.

Why it's built this way

An explicit graph with typed state, so each step can be tested and traced on its own.

Connections

UI

Designed for verifying, not just reading

The interface assumes people will check the answer, and makes that effortless. Citations are inline, sources preview on hover, and the assistant is honest when it can't find something.

Opening a citation jumps to the exact passage in the source document, highlighted.
The admin view of unanswered questions on desktop, next to the assistant on mobile.
  • Citations inline

    Every claim links to the passage it came from; hovering a citation previews the source.

  • Honest about gaps

    If nothing relevant is found in documents you can access, it says so instead of guessing.

  • Feedback in one click

    Thumbs up or down on every answer, with an optional note that reaches knowledge managers.

Development

Evaluation before features

RAG systems fail quietly, so the build started with a way to measure answers. Every change to chunking, retrieval or prompts runs against the same graded questions before it ships.

  • Graded question set

    Real questions with answers graded by knowledge managers, run on every change to the pipeline.

  • Typed graph state

    The LangGraph flow has typed state and explicit nodes, so each step can be unit-tested.

  • Connector contract

    One interface for every source (list, fetch, permissions, changes), tested against recorded responses.

  • Tracing

    Each answer records the rewritten query, retrieved passages and model output for debugging.

Stack

Frontend
Next.jsTypeScript
AI
LangGraphLLM
Data
pgvectorRedis
Sources
SharePointGoogle DriveConfluence

AI

Retrieval that respects permissions

Documents are chunked by structure, embedded, and stored in pgvector with their access groups. At question time the graph rewrites the question, retrieves with the person's groups as a hard filter, re-ranks, and generates an answer that must cite what it used.

A cited answer

Sample data

Waiting for a questionA consultant asks a policy question. The assistant searches only what they can access and cites each claim.
  1. 1

    Understand the question

    Follow-ups are rewritten into standalone queries, and the graph decides whether retrieval is needed.

  2. 2

    Retrieve with a permission filter

    Vector and keyword search run with the person's access groups inside the query, not after it.

  3. 3

    Re-rank and assemble

    The most relevant passages are ordered and trimmed to fit, keeping their source links.

  4. 4

    Answer or abstain

    The model answers only from those passages, cites them, and says so when they don't cover the question.

Guardrails

  • Access filtering happens inside the database query, so restricted text never reaches the model.
  • Answers without supporting passages are replaced by an honest 'I couldn't find this'.
  • Every answer stores its retrieved passages, so a bad answer can be traced and fixed.

Integration

Three sources, one permission model

Each source expresses access differently. Connectors translate SharePoint permissions, Drive sharing and Confluence restrictions into one model of access groups that retrieval can filter on.

  • SharePoint

    Microsoft Graph API

    Sites and libraries with item-level permissions; delta queries fetch only changes.

    Inbound
  • Google Drive

    Drive API

    Shared drives with file permissions; the changes feed keeps the index current.

    Inbound
  • Confluence

    REST API

    Spaces and pages with space and page restrictions resolved to groups.

    Inbound

Performance

Quick answers without cutting corners

Speed matters in a chat interface, but not at the cost of permissions or grounding. The approach keeps every safety step and makes each one cheap.

  • Incremental indexing

    Connectors use delta and change feeds, so only new or edited documents are re-chunked and re-embedded.

  • Indexed, filtered search

    An approximate-nearest-neighbour index in pgvector plus an index on access groups keeps filtered search fast.

  • Streaming responses

    Answers stream as they're generated, so people start reading straight away.

  • Caching repeat questions

    Retrieval results for frequent questions are cached in Redis per permission set and invalidated when sources change.

  • Small, focused context

    Re-ranking trims the context to the passages that matter, which keeps answers faster and cheaper to generate.

Outcome

What's in place so far

In build: outcomes below describe what works today

The assistant is in build. Retrieval, citations and the feedback loop work end to end against the connected sources, and the unanswered-questions view is ready for knowledge managers. Launch follows once permission sync has been verified source by source.

What's next

After launch, answer feedback and unanswered questions become the main signals for what to improve, both in the assistant and in the documentation.

In place so far

  • Answers link back to the source document so staff can verify them

Outcomes are described qualitatively. Client figures stay with the client.

Have a similar project?

Whether it's an AI feature that needs to be trustworthy or a system you need to integrate with, tell me what you're building. I reply within one business day with questions and a suggested first step.