Skip to content

Knowledge

HangRui Index

Your documents, your machine, answers with sources.

"It has to know our company's material" and "the material cannot leave the building" are usually two requests from the same person. Index keeps the knowledge in an index you hold and can rebuild at any time, rather than in a model's weights. The model only reads the few pieces that were chosen, so it can be one racked in your server room or an outside API, and switching is one line of configuration.

Available on
  • Web

The assistant on this site, one real run

Recorded September 2, 2026; the counts are from that run

Can HangRui Index run with a local model? Does our data leave the building?
Ask

Can HangRui Index run with a local model? Does our data leave the building?

Yes. HangRui Index can run with a local model on your own hardware, so the documents never leave your environment 1. If you use an outside API instead, only the few selected pieces are sent — never the whole document 1.

How it answered

  1. Retrieve5/37The site is cut into pieces and every one is scored against your question. Only the highest scoring few are put in the prompt; the model never sees the rest.
  2. Answer1/5The model may answer from those pieces and nothing else. What it knows from anywhere else does not count here.
  3. Verify1/1Every citation it wrote is matched back against what was actually sent. A number it invented is left as plain text instead of becoming a link.
  • Answered from
  • [1]HangRui Index100%
  • [2]Who else touches your data68%
  • [3]HangRui Pages65%

Every number in the answer links to the page that sentence came from.

Ask one yourself

What it does

Documents stay in the building

Cutting, scoring and selection all happen in your environment. With a local model, no text leaves your machine at any stage. With an outside API, what leaves is the few pieces that were chosen, never the whole document.

Local or cloud, one line

The model is the only stage that knows who the supplier is. Run an open-weight model such as Llama or Gemma on your own GPU, or point it at an API such as Anthropic. Changing is an environment variable, not a rewrite.

Every citation is verified

Each number in an answer is matched back against the pieces that were actually sent. Matches become links; anything else stays plain text and is flagged, so an invented source cannot pass as a real one.

Editing a document is the update

Change the document, rebuild the index, and the next answer is current. No training run and no regression suite, because the knowledge was never learned into the weights.

How it works

Point it at your documents

PDFs, Word files, a help centre, a ticketing system, or a shared folder. It cuts along the structure the documents already have.

Choose a model

An open-weight model in your own server room, or an outside API. Both take the same path, and the same numbers are printed.

Read how it answered

Under every answer, three counts: pieces read, pieces sent, citations verified.

Local or cloud, the same path.

Only the model stage knows the supplier. This is what each choice buys, and what it costs.

A model on your own hardware

  • Neither the documents nor the questions leave your machine
  • Prompt length is your own GPU time, so only the most relevant pieces are sent
  • No per-call fee, so unit cost falls as volume grows
  • Someone has to look after the hardware and the model version

An outside API

  • What leaves is the few chosen pieces, never the whole document
  • Access to the largest models, which do not fit in a workstation
  • No GPU to run, and you pay by use
  • Content passes through the supplier, and the privacy policy has to say so

The retrieval, the prompt and the verification are the same code on both sides, and so are the three counts printed on the screen.

The other eight

Have something in mind?

Tell us what you are building and we will tell you honestly whether we are the right studio for it.