Ten milliseconds, because there is nothing to think about
The hotel’s own answers are written before the call, so a common question never reaches a model. Every other agent — ours included, on the harder turns — has to wait for one.
Technology
Most voice agents send every sentence to a large model and wait. Cilicia keeps the hotel’s answers ready, learns the caller while the phone is still ringing, and reserves the confirmation number before the database has finished writing it. This page shows how — and how that compares with PolyAI.
Measured on the demo line · Crowne Plaza Englewood
A question the hotel has answered before leaves in 10 ms.≈440 ms when the model has to think — about half a second in your ear either way.
Architecture
Three moments from one call, with a stopwatch. What you say, what you hear, and how long you waited in between — and, in one line, why it was that short.
Crowne Plaza Englewood
Cilicia AI · 00:00
Loaded while ringing
Photo: randomuser.me · a sample caller, not a real guest
0.00 s
you wait for the greeting — nothing to look up once you speak
While your phone is still ringing, Cilicia has already looked up who you are, the weather at the hotel and the date. The greeting knows you — nothing to fetch once you start talking.
Built different
PolyAI is an excellent company: nine years in, more than $200 million raised, valued at $750 million, serving ten industries, with models of its own. Cilicia answered its first call in March 2026. Their model is faster than ours. We are faster than their model — on the questions a hotel is actually asked all day, because we answer those without one.
Cilicia 2026 · Wyoming
PolyAI 2017 · London · as reported
The hotel’s own answers are written before the call, so a common question never reaches a model. Every other agent — ours included, on the harder turns — has to wait for one.
Guest, weather and date are loaded while the phone rings. Others fetch them after you start talking — and you hear the pause.
A platform for ten industries has to be general. Cilicia only knows hotels — the pillow menu, the late check-out, the explicit yes.
Staff type “the pool is closed Monday mornings” in the panel and it is live on the next call, in all eight languages. No studio, no ticket.
Nothing upfront and no per-minute meter: ten percent of the revenue booked over the phone. If it does not book, it does not bill.
Questions it had to think about are queued for staff; approved answers become instant answers. The second month is faster than the first.
PolyAI figures as reported by SiliconANGLE and others, December 2025 (Series D, $86M at a $750M valuation; more than $200M raised). Not an affiliation or endorsement.
When the hotel’s facts change, Cilicia writes the answers to a hundred likely questions in eight languages and keeps them ready. A matching question is answered from memory in about 10 milliseconds; only the rest goes to the model.
The caller ID arrives before the first ring ends. In those seconds Cilicia looks up the guest, the weather at the hotel and the date, and hands them to the agent — so the first sentence can already say “welcome back”.
The confirmation number is reserved and read out at once; writing the record and texting the payment link finish in the background. No “one moment while I confirm that”.
Generation starts during the guest’s pause, the voice begins on the first sentence, and the model behind it is the fastest we can find that still books a room correctly.
One turn, to scale
The model lane is the honest number — it includes detecting that the guest has stopped, the model’s first word, finishing the first sentence and the voice. The memory lane skips the model entirely, which is why “is parking free?” comes back in half a second.
Common question · memory lane
0.00 s
Anything else · model lane
0.00 s
Both lanes on the same clock, slowed down 4×.
Only the middle bar is ours to measure: 10 ms from memory, 440 ms to the model’s first word (medians, demo line, September 2026). Silence detection, voice and the phone network are typical values for this carrier and move call to call. Booking turns add one tool call; the confirmation number is spoken from the first model reply.
Compared
Technical only — what each system does and how long it takes. Both sides of the speed rows measure the same thing, the way PolyAI defines it: from the request reaching the brain to the first text the voice can start speaking.
Where we are quicker
Parking, breakfast, check-out, pets — the questions a front desk answers all day. Cilicia has written those answers in advance, so the reply leaves in ≈10 ms with no model in the loop. PolyAI’s own published figure for its fastest model is under 300 ms, because every turn goes through one. On those questions that is roughly 30× less waiting.
Where they are quicker
When a turn genuinely needs thinking, we call a general model and wait ≈440 ms for its first word. PolyAI runs its own model on its own GPUs and publishes under 300 ms. That is faster than us, and it is their engineering, not marketing. We would rather you read it here than find it later.
| Cilicia | PolyAI as published on poly.ai | Edge | |
|---|---|---|---|
| A question the hotel has answered before | ≈10 ms. The answer was written before the call, in eight languages, and no model is involved (measured, Sep 2026) | Every turn goes through a model; the published figure for their fastest, Dialog-RSN-1, is “under 300 milliseconds” | Cilicia |
| A turn that needs thinking | ≈440 ms to the first word — median on the demo line, GPT-4.1 on OpenAI’s priority tier (measured) | Sub-300 ms on their own Dialog-RSN-1, served on A100 GPUs (published) | PolyAI |
| What produces the answer | Memory first: the hotel’s own facts become ready answers; only what misses reaches a general model | Proprietary models — Raven, trained on 1B+ enterprise conversations, and the audio-native Dialog-RSN-1 (published) | Even |
| Between calls | The 6,000-token hotel prompt is kept alive in OpenAI’s cache; 96–99% of it is served from cache inside a call (measured) | Own models on own GPUs; nothing published on caching between calls | Cilicia |
| What it knows when it picks up | Caller, weather and today’s date are loaded before the first ring ends — nothing to look up once the guest speaks | Recognises loyal guests through a PMS integration (published) | Cilicia |
| Hearing the caller out | Patient endpointing; hesitations — “eee”, “hmm” — never end a turn and never count as an interruption | Dialog-RSN-1 is audio-native: it reads pauses and tone from the audio itself rather than from a transcript (published) | Even |
| Languages | Eight — English, Turkish, Spanish, German, French, Arabic, Korean, Hindi — switching mid-call | Broad multilingual coverage across ten industries; poly.ai publishes no count | PolyAI |
| Certifications | Consent announced on every call, full transcripts in the panel; formal certifications in progress | SOC 2, HIPAA, GDPR, PCI DSS as standard (published) | PolyAI |
Scoring is ours: one point per row to the side with the stronger published fact for a 50–300 room hotel; “even” where both do it or the rows are different by design. PolyAI statements are quoted from poly.ai — the sub-300 ms figure, the A100 detail and the audio-native description from their Dialog-RSN-1 post, the certifications and Raven from their home page — read September 2026; where we found no public figure we say so rather than guess. Cilicia figures are medians from our own demo line in September 2026 and will move with the carrier, the guest’s phone and the model we run. Neither number is what the guest hears end to end: both leave out the silence detector, the voice and the phone network, which add roughly another 400 ms on each side. PolyAI is a trademark of its owner; no affiliation, no endorsement.
Try the architecture, not the slides
The first answer comes from memory, the second from the model, and the confirmation number arrives before the text does. Then watch all three land in the panel.
Photos: Unsplash.