Every enterprise AI vendor demo hits the same beats: a clean interface, a confident answer, a login page that looks secure. None of that tells you what happens to your data once the demo ends. The questions that actually matter get asked later, if they get asked at all. Usually they come from IT or legal, usually after the business side has already fallen in love with the product.
Ask them earlier instead. The questions below are not a compliance checklist to file away; they are the questions that separate a real AI intelligence layer from a well-designed demo with a login page bolted on. They are also the same questions an internal build would have to answer for itself, which is worth remembering if buying versus building is still an open question for you. A vendor with a real answer will give you a sentence. A vendor without one will give you a paragraph.
Where does your data go, and is it copied or queried?
Start with the most basic question: does your data ever leave your systems? Some platforms copy data out of your ERP, CRM, or database into their own store to make it faster to search. Others query your systems live, in place, and never hold a persistent copy.
Neither answer is automatically disqualifying, but the difference matters. A copy is a second place your data now lives: a second attack surface, a second thing to secure, a second place it can go stale. A live query means the platform is only ever as exposed as the connection itself, and a revoked credential shuts the door completely.
A good answer names exactly what leaves your environment and why: "we cache query results for 15 minutes to keep dashboards fast; nothing else is stored." A red flag is "your data is processed securely in our platform," which describes nothing.
Does our data train your models?
This is the question with the most consequences and, too often, the vaguest answers. If a vendor uses your data — including your prompts and the documents you connect — to train or fine-tune models, that data can, in principle, resurface in ways you did not intend, for customers you did not choose.
This should be a yes-or-no question with a written answer, not a verbal reassurance in a sales call. Ask specifically about their own models and about any third-party model providers they route through, since "we don't train on your data" can quietly mean "but our upstream provider might."
A good answer is unambiguous and in writing: your data is not used to train any model, full stop, or here is exactly which narrow signal is used and how to opt out. A red flag is "we take privacy seriously" without a direct answer to the actual question.
Whose permissions do the answers respect?
This is the one buyers skip most often, because it only shows up after rollout. An AI system that connects to your ERP or CRM needs credentials to do that. The question is whether it queries as you, respecting exactly what your account can see, or as one shared service account that every user's questions run through.
The difference is invisible in a demo and very visible six months in, when a junior staff member asks a question and gets back a number they were never supposed to see, because the AI answered with the service account's permissions, not theirs. Per-user enforcement is the difference between a governed platform and a shared password with a chat window in front of it.
A good answer describes how the platform maps your existing role and permission structure onto every individual answer, not just onto system access. A red flag is any version of "the AI has one connection to your systems," offered as if that settles the question.
Who can see that someone asked, and what they asked?
Once AI can see sensitive data, the question of who asked what — and when — stops being optional. An audit trail is not about distrust of your own employees; it is about being able to answer, after the fact, exactly what data an AI touched and on whose behalf.
Ask whether the log covers the question, the answer, the data sources touched, and the identity of the asker. Ask too whether it is something you can export and query yourself, or something you have to request from support.
A good answer is a log you can query directly, covering who, what, when, and which systems were touched. A red flag is "we can look into it for you," a manual, vendor-mediated process instead of a trail you own.
How long do you keep it, and how is it deleted?
Retention is where good intentions quietly rot. A platform that stores query history, cached results, or conversation logs indefinitely has created a growing pile of your business data outside your control.
Ask for a specific retention period, not a range, and ask what happens when you delete something: does it disappear from backups too, and on what timeline? Also ask what happens to everything if you leave the platform entirely: is there a clean, documented deletion process, or does your data simply linger because nobody asked.
A good answer names a specific window, a specific deletion mechanism, and a specific timeline for backups. A red flag is "we retain data as needed for service quality," which retains the right to keep everything forever.
What about WhatsApp and other chat interfaces?
WhatsApp is a genuinely useful interface for enterprise AI precisely because people already live in it. That is exactly why it needs the same scrutiny as any other access point, not less. A question typed into a chat app still needs to resolve to a specific, verified identity before it touches your systems, and the answer that comes back still needs to respect that identity's permissions.
The convenience of chat should never come at the cost of the permission model you'd insist on for a dashboard login. If a vendor cannot explain how a WhatsApp number is tied to a verified employee identity with a defined scope of access, that convenience is a liability wearing a familiar interface.
A good answer walks through identity verification and permission enforcement for chat exactly as it would for a web login. A red flag is treating the channel as an afterthought ("we just connect to WhatsApp Business API") with no mention of identity or scope.
More questions worth asking before you sign
How is your data isolated from the vendor's other customers: a genuinely separate environment, or shared infrastructure with logical boundaries that depend on code being correct every time? What is the vendor's incident posture — not a promise of perfection, but how and how fast you would actually be told if something went wrong? And what happens to your data and operational continuity if you stop using the platform: a documented exit path, or starting from zero with a new vendor?
Four more are less glamorous and just as decisive. Is your data encrypted both at rest and in transit, and who holds the keys? Which sub-processors and third parties touch your data, does the vendor disclose the full list, and are you notified when it changes? Where does your data physically reside, and can the vendor keep it in a specific region if data residency or Indonesian localization requirements apply to you? And how does the platform handle identity at scale: does it support SSO and SCIM provisioning, so that access is granted and, more importantly, revoked automatically when someone's role changes or they leave?
None of these are exotic questions. A serious vendor has already answered them for other customers.
The checklist
| Question | What a good answer sounds like | Red flag |
|---|---|---|
| Does data leave our systems? | Names exactly what data moves and why | "Processed securely in our platform" |
| Copied or queried live? | Specific mechanism and retention window | Vague reassurance, no mechanism |
| Used to train models? | Written "no," or a named, opt-out-able exception | "We take privacy seriously" |
| Whose permissions apply? | Per-user enforcement on every answer | "One connection to your systems" |
| Audit trail? | Exportable, queryable log of who/what/when | "We can look into it for you" |
| Retention & deletion? | Specific window, specific deletion mechanism | "As needed for service quality" |
| WhatsApp / chat access? | Identity verification tied to permission scope | "We just connect to the API" |
| Multi-tenant isolation? | Named isolation model, not just "secure" | Deflects to a certification logo |
| Incident posture? | Defined notification timeline | No commitment on how you'd be told |
| Exit path? | Documented export and deletion on departure | No answer, or "talk to your rep" |
| Encryption at rest & in transit? | Both, with a clear answer on who holds the keys | "Everything is encrypted," no specifics |
| Sub-processors disclosed? | Full list, with notification when it changes | No list, or "industry-standard providers" |
| Data residency / localization? | Named region, honored on request | "Our cloud is global," no control offered |
| SSO & SCIM provisioning? | Automated provisioning and de-provisioning | Manual user management, no SCIM |
None of this replaces your own legal and security review; it makes that review faster, because a vendor's answers here tell you whether the rest of the diligence will be straightforward or a fight. It also assumes your own data is in a state worth connecting in the first place, so if you are not sure, enterprise data readiness is the prior question. In Indonesia specifically, the Personal Data Protection Law (Law No. 27 of 2022) has made this a board-level topic rather than an IT checkbox: the company deploying the AI is the data controller, and it carries responsibility for how personal data is handled regardless of which vendor built the platform.
Where Nalar fits
Nalar is built around the assumption that these questions get asked, and answered specifically. The platform connects to the systems you already run with live, governed access rather than copying your data elsewhere, enforces your existing per-user permissions on every answer, and keeps every question and its data sources traceable, including questions asked over WhatsApp.
We are not going to make certification claims here; that is exactly the kind of thing this article warns you to look past. What we would rather do is answer these ten questions directly for your specific setup. The interactive demo shows how the permission and access model actually behaves across a realistic enterprise. And if you are not yet sure where your own data stands before that conversation, BARI — our AI-readiness diagnostic — will tell you honestly, including when the answer is "not yet."
Frequently asked questions
- What is the single most important data-security question to ask an AI vendor?
- Whether your data is used to train their models. It is a yes-or-no question with real consequences, and a vendor who answers it with a paragraph instead of a sentence is telling you something.
- Should every enterprise AI evaluation include a security review?
- Yes, before a pilot ever touches real data. A structured review (data flow, permissions, retention, audit) takes a fraction of the time a breach or a regulatory finding does, and it belongs in procurement, not left until after rollout.
- Does Indonesia's PDP Law change what enterprises should ask vendors?
- It raises the stakes without changing the underlying questions. Data governance for AI vendors was always a board-level topic — the law makes it explicit that the company, not the vendor, bears responsibility for how personal data is handled.
- What is a red flag in how a vendor answers these questions?
- Vagueness dressed as reassurance: 'your data is safe with us,' 'we take security seriously,' or a pivot to a certification logo instead of a direct answer about data flow, training use, or permissions. A governed platform can answer each question in one specific sentence.