AI phone assistant
The assistant picks up, speaks the mandatory notice and files a task. This page walks through a call in order, including the places where it deliberately promises nothing.
Last updated: 2026-08-20
Plenty of vendors now promise that a voice will pick up. The questions that decide whether you keep using it come later in the call: what does the caller hear first? What happens to what was said? And what does the system do when it cannot understand the caller, or cannot speak itself?
This page answers those in the order they happen on a call.
Before the AI speaks
A call reaches your property through a number booked to your property. That mapping is the first check, and it is strict: a call arriving on a number that belongs to no property is refused rather than landing with whichever tenant happens to be nearby.
Three further checks follow, in this order.
Country of origin. You can decide which countries calls are accepted from. If a list is configured and the caller's number cannot be placed, the call is refused. Somebody who configured a restriction would rather lose a call than lose the restriction.
Daily limit. There is a per-property daily cost ceiling that you set. Once it is reached, the assistant takes no new calls that day.
Opening hours. With no hours configured the assistant answers at any time. With hours, your property's timezone applies, overnight windows across midnight included, and a single date overrides the weekly rule.
All three are built so that a missing or malformed configuration lets the call through. Only a valid setting that genuinely says "no" refuses. A typo in the admin must never send a guest to a dead line.
The first sentence is a legal obligation
Before the assistant listens, it speaks. The opening names your property, says that a digital AI assistant is speaking, notes that the conversation is processed, gives the address of the privacy notice, and says the caller can hang up.
That is not a pleasantry. It is three duties in one utterance:
- Article 50 of the AI Act. A person must know they are talking to an AI. So the voice introduces itself as a digital assistant, audibly, in the first sentence.
- § 201 of the German Criminal Code. We treat real-time transcription as a recording of the non-publicly spoken word, whichever recognition service is configured. Which is why the objection has to be possible during the notice, not after it.
- Article 13 GDPR. A short spoken notice cannot discharge the information duties in full. So it names the address where the complete privacy notice lives. Without that reference the notice would be incomplete.
Two properties of this opening are technical rather than cosmetic.
It is not interruptible. Anyone who could skip it would not have heard it, and consent nobody heard is not consent.
And it runs before the caller's audio reaches speech recognition. The order is the whole point: inform first, process after.
The wording is versioned. Every playback is logged with the version it was spoken under, so it stays possible to work out later which text was agreed to. If the text changes, the version moves with it.
Speaking over the notice is not the same as being ignored
An uninterruptible opening has a side effect you have to have seen live once.
On 13 August 2026 a caller spoke five seconds into the opening. The framework discards what is said during uninterruptible speech: no reply, and no entry in the conversation either. The sentence was simply gone. The caller sat through the rest of the notice and was then asked how we could help, having just said so.
The notice stays uninterruptible, because it has to be. What the caller says over it is now buffered and answered the moment the notice ends. Both requirements can be met, just not on the default setting.
Recording, objection, and what is kept
A caller can object to audio recording while the call is running. Two ways: key 2 on the phone, or simply saying so.
Both reach the same result: audio recording is switched off, the transcript is kept. The objection can be raised more than once without harm, and it is logged.
Why the transcript stays rather than going with it: it is what makes the request get handled at all, and it is what an access request by the same caller later has to be answered from.
Listening is the harder half
Speaking is a solved problem. Listening is where a phone assistant gets embarrassing, and the three rules below all came out of real calls.
The room in the background is not the guest. On a call in July 2026 the words "privat" and "Rekord" were committed as guest turns and fed to the model. Noise cancellation was running and was not the gap: it removes steady noise, fans and hiss. A television in the background sounds like speech and passes straight through. The gap sits one layer lower, in recognition itself, and is closed there on the recognition's own confidence rather than on word count. "Double room", "August" or a single spelled letter are one-word answers too, and dropping a real answer is the worse failure.
The language switches mid-call, but not on request. The text of the answer is authoritative; the recognition service's language label may only override it under conditions. Otherwise one mislabeled sentence flips a whole call into another language. There is more on multilingual work, including your team's side of it, on Multilingual.
"Are you still there?" does not fire while the guest is waiting on us. The question makes sense when a caller has walked away and is embarrassing when they are waiting for an offer. So the trigger reads the assistant's own lifecycle state and not merely whether a tool is running: on a call in July 2026 the availability lookup answered in milliseconds while the model and the speech output took seconds, and those seconds were exactly the unguarded ones.
What gets said passes through a chain
Between the model and the voice there is exactly one place where text is changed, and it is a chain of single steps, sentence by sentence: echoed instructions and list markers are stripped, then your property's blocked patterns apply, then the pronunciation overrides you can configure, then whatever has to be spelled out is spelled out.
The last part is the one you hear immediately on a phone. An email address is read back letter by letter, a web address is spoken rather than read out as a string, and runs of digits are said one at a time in the language currently in use. A Wi-Fi password read out as a number is useless on the phone.
What can be done on a call
Booking. Check, hold, payment link, confirm. The sequence, the 15-minute hold window and the limits are covered at length on Direct bookings and are not repeated here.
Sending a link to a mobile. A form or a payment link reaches the caller in a fixed order: WhatsApp, then SMS, and if neither works, a ticket for your team. A send that fails does not count as a delivery. And when the call has no usable number, the assistant asks for one once rather than filing a ticket straight away. A banner on a screen exists only where there is a screen, which means the browser demo and not the telephone.
Handing over. Where the assistant cannot get further, it files an urgent task and marks the conversation as escalated. What happens next is described on AI ticketing.
When the AI cannot speak
On a call in July 2026 a caller sat through two minutes of silence while every synthesis attempt failed. There was a fallback voice, it did not engage, and the system kept generating replies into a dead line.
Today, if speech output fails completely, no announcement can be synthesised any more. So a pre-rendered recording is held ready for that case: the caller hears, in their own language, that there is a fault, and the call is ended cleanly once the message has played out.
A fault you can hear beats an open line where nothing happens.
After the call
The transcript is stored as a message history, in the same list as chat, WhatsApp and email. There is no second place for "the phone calls".
It is written during the call, not only at the end. The reason is an incident: a seven-minute call in July 2026 ended with not one sentence stored, because the teardown routine never ran and the whole transcript died with the buffer. The call showed up empty in the admin, and it would have been missing from an access request under Article 15 GDPR.
What happened on the call is additionally in the audit log: accepted, refused and why, notice played, tool executed or refused, call ended. That log can only be written to, never changed, including by us.
Access and erasure on the phone
There is a tool that lets a caller exercise their rights during the call: access to what is stored against their number, or erasure.
Three limits come with it. The tool is not on by default, it is enabled per property. Access returns counts to the model and never content, because an access record has no business inside a language model. And erasure is final, which is why it sits behind the confirmation gate: it runs only after a separate, explicit confirmation from the caller.
Where processing happens
A call's components are assembled at exactly one place in the code, and EU regions are set there: the language model through Vertex AI in an EU region, speech output through its provider's EU endpoint, speech recognition on a European region. Turn detection runs inside the process itself and calls nothing outside.
That it sits in one place is the actual point: there is a file you read to know what a call touches, instead of a search across the whole codebase.
Which providers sit behind those, and the terms they are engaged on, belong in the data processing agreement rather than on a product page. How to obtain it is on Data processing agreement; what is processed in general is in the privacy notice.
Limits
No live transfer to a person. A handover files an urgent task and marks the conversation, but it does not bridge the caller through to somebody. The wiring for that is not in place on today's telephony path. That is the more honest answer than a promise that shows itself up at the decisive moment.
Outside opening hours the caller hears nothing. The call is refused, but the spoken announcement for it is not played on today's telephony path. Anyone setting opening hours should know that the guest does not hear "we are not available right now" first, only the end of the call.
The opening is German to begin with. The caller's language is recognised from their first sentence of their own. For a property in the German-speaking market that is the right default, but it is a default.
No speed promise. This page carries no response time, no number of seconds and no percentage. Not out of modesty: there is no measured, dated series for the telephony path that would carry such a figure, and this website has already published one number that was wrong by a large multiple and had to be removed. Where the figures that are here come from is set out in the methodology.
The difference
Most pages about AI on the phone sell the moment the voice picks up. That is the easy part.
The parts that decide whether a property is still using the system four weeks later are quieter: that the mandatory notice is actually heard, that a television in the background does not trigger a booking, that an outage is audible rather than silent, and that the day after it is possible to read what was said.
Which is why so much of this page is about calls that went wrong. They are the reason the rules look the way they do.