Support
HangRui Desk
It answers on LINE, and it says when it does not know.
Most support bots are confident and wrong, which costs more trust than they save in tickets. Desk answers only from documents you gave it, shows the source, and escalates rather than guessing — on LINE, on Telegram, and in a widget on your own site, with the same answer and the same counts on all three. The whole pipeline also runs on hardware you own, which for a lot of businesses is the condition for letting it near customer conversations at all.
- Web
Ask this site.
Answers come from a model reading this site’s own pages. It is not a person, and it knows nothing that is not published here.
What it does
Grounded in your docs, and it stops when they run out
It answers from your help centre, policies and past tickets. When nothing in that material matches the question, it does not call the model at all — it says so and offers a person.
On the channels people already have open
LINE, Telegram, and one line of script on your own site. One brain behind all three, so the same question gets the same answer wherever it was asked.
Every citation is checked, not just printed
Each number in the answer is matched back against the passages actually sent to the model. Ones that match become links; invented ones stay as plain text and are marked as unverified.
It can run on a machine in your shop
The same pipeline with an open-weight model on your own GPU. Customer messages and your documents stay on hardware you own, and there is no per-question bill.
How it works
Point it at your material
Help centre, returns policy, shipping table, opening hours, resolved tickets. PDF, Markdown or a URL.
Connect the channel
Your LINE official account, a Telegram bot, or one line of script on the site you already have.
Set the handover rules, then read what it could not answer
Which topics always reach a person — and once a week, the questions it had no document for. That list is the next document to write.
Channels
The answer has to turn up where the customer already is.
One brain, three adapters. The same question gets the same answer on all three, with the same three counts under it: passages read, passages sent, citations checked.
LINE
A reply card carrying the answer, one button per source that opens the page it came from, and a button that hands the conversation to a person. A loading indicator runs while it reads.
Replying to a customer's message is free and pushing a message later is metered, so escalation notices go by email rather than push.
Telegram
The same answer, written out as it is produced: the message is edited in place, so the text grows the way it does on a web page. Sources are inline buttons and /human reaches a person.
Edits are throttled to one every 1.5 seconds, because an edit that does not change the text comes back as an API error.
On your own site
One line of script. It opens an iframe, so your stylesheet cannot reach in and ours cannot leak out, and the panel reports its own height back so it never opens half empty.
The same panel as the assistant on this page: the counts, the sources and the runners-up are all printed, because the mechanism is the product.
Measured 2026-09-02: the same question on LINE and on Telegram read 28 passages, sent 5, and verified 1 citation on each. Identical across channels is the acceptance condition, not a bonus.
Measured
What happens when the harness comes off.
Same model, same 28-passage corpus, same 30 questions. The only thing removed is the code around the model: picking the passages, stopping when none of them match, and checking the citations afterwards.
| Measure | With the harness | Without it |
|---|---|---|
| Passages sent to the model | 3 (240 characters) | 28 (2,076 characters) |
| Input tokens per question | 497 | 2,768 |
| Cost per question, one assumed rate card | $0.000647 | $0.002959 |
| Questions answered acceptably | 30 / 30 | 23 / 30 |
| When the answer is not in the documents | The model is never called | A fluent answer, with a citation that does not match it |
| Claimed a source that could not be verified | 0% | 20% |
Two runs, both 2026-09-02 on our own hardware. Passage and character counts come from the qwen2.5:7b pass (30 questions × 3); tokens, cost and the verification rates from the llama3.1:8b pass (30 questions × 1). The rate card is an assumption used to show the arithmetic, not a quote. Twenty-eight passages is a small corpus and none of this transfers to your documents until it has been re-run on them.
Our machines, or one of yours.
Only one box in the pipeline knows which model it is talking to, so the same Desk runs both ways. What changes is where the conversation goes.
On a machine in your shop
- Customer messages and your documents stay on hardware you own
- No per-question billing — the cost is the machine and whoever looks after it
- The model has to stay resident: a 14B model's first call measured 20.7 s here, most of a reply budget
- The public address is a Cloudflare tunnel, which costs nothing; the domain is the bill
On ours
- Answering the same day, with nothing to install
- Reaches models too large for a single workstation
- Customer messages pass through a model provider, which your privacy page has to say
- Billed by usage, which is why the passage count above is the number to watch
Retrieval, prompt and verification are the same code either way. Across four local models from 4B to 14B the retrieval numbers came out identical: the model changes how the answer reads, not whether the right document was found.
The other eight
- KnowledgeHangRui IndexYour documents, your machine, answers with sources.
- WebHangRui PagesA site that ships, not a mockup.
- AppHangRui ScreensFrom a screen list to something you can install.
- LyricsHangRui LyricsThe rhyme is checked, not promised.
- Wedding filmHangRui VowsYour story, one scene at a time.
- MusicHangRui ScoreArrangements you can still take apart.
- ImageHangRui StillsA look you can repeat.
- VideoHangRui SequenceA rough cut before the budget.
Have something in mind?
Tell us what you are building and we will tell you honestly whether we are the right studio for it.