AI Data Readiness Checklist: 5 Questions Before You Start
Five questions to answer before an AI project starts, so week one is spent building rather than searching for spreadsheets.
Key takeaways
- Most delays in AI projects come from finding, cleaning and getting access to data, not from the model.
- Five questions cover the basics: where the data lives, whether you can grant read access, how far back it goes, whether it holds personal data, and who can confirm a right answer.
- If you can answer three of the five, the project can start; the rest can be worked through in week one.
- Personal data should be masked or excluded unless the AI genuinely needs it, in line with the data protection law where you operate.
Before an AI project starts, answer five questions: where does the data live and who owns it, can you grant read access in the first week, how far back does it go and is it consistent, does it contain personal data that needs masking, and who can confirm when an AI answer is right. If you can answer three of the five, the project can start.
Most delays in AI projects are not about models. They are about finding, cleaning and getting access to the data the model needs. A team can lose the first week of a sprint waiting for a database password or discovering that “the sales data” is actually six spreadsheets with different column names. This checklist exists so that week one is spent building.
The checklist
- Where does the data live today, and who owns it?
- Can we get read access in the first week?
- How far back does it go, and is the format consistent?
- Does it contain personal data that needs masking?
- Who can confirm when an AI answer is right?
If you can answer three of the five, the project can start. We work through the rest together in week one.
Below is what each question is really asking, and what a good enough answer looks like.
1. Where does the data live, and who owns it?
Data in SMEs is rarely in one place. It is usually split across an accounting system, a few spreadsheets, a shared drive, email attachments and a WhatsApp group or two. That is normal, and it is not a reason to wait.
What we need is a list. For each source, write down what it holds, where it is, and the name of the person who can say yes to sharing it. Ownership matters because access requests stall when nobody feels able to approve them.
A good enough answer
- A list of sources, even if rough
- One named owner per source
- A note of which sources the first version actually needs
You rarely need every source. A first version that tags products might need only the product catalogue. A first version that answers customer questions might need only the FAQ document and the price list.
2. Can we get read access in the first week?
Read access means the ability to look at data without being able to change it. It is safer for you and enough for most of the build.
Access is the most common hidden delay. Requests go to IT, IT is waiting on a vendor, the vendor needs a form signed. Starting that process before the project begins saves days.
Ways to provide access
| Option | When it works well |
|---|---|
| Read-only database user | Data sits in a database you control |
| Scheduled export to a shared folder | The system has no direct access, but can export files |
| API key with read scope | The system offers an API, such as many cloud tools |
| One-off export | Historic data for testing, with live access to follow |
Any of these is fine to start. A one-off export is often the quickest way to begin while proper access is arranged.
3. How far back does it go, and is it consistent?
History helps the AI learn patterns and gives us real examples to test against. Consistency matters just as much. Three years of data where the product code format changed twice, and where the “status” column means different things in different years, needs cleaning before it is useful.
What to check
- When the data starts, and whether there are gaps
- Whether the same field means the same thing over time
- Whether codes, names and units are written consistently
- How many records are blank, duplicated or obviously wrong
You do not need to fix any of this before starting. You only need to know roughly how it looks, so the plan accounts for the cleaning.
Take a feature like the AI defect detection in our field inspection project in Kerteh, Malaysia. For work of that kind, past inspection photos and findings are the raw material, so knowing early which records are usable shapes what a first version can realistically do.
4. Does it contain personal data that needs masking?
Customer names, phone numbers, addresses, identity numbers and health information all count as personal data. The rule we follow is simple: share only what the AI genuinely needs, and mask or remove the rest before it leaves your systems.
If the AI does not need a person’s name to do its job, it should never see one.
Know which law applies
Each market has its own law:
- Singapore: the Personal Data Protection Act. The Personal Data Protection Commission publishes guides and advisory notes.
- Malaysia: the Personal Data Protection Act 2010, overseen by the Personal Data Protection Department.
- Indonesia: Law No. 27 of 2022 on Personal Data Protection (UU PDP).
This is not legal advice, and your obligations depend on your business. If you handle sensitive data, such as patient records, involve your data protection officer early. Healthcare work like our patient app for 11 hospitals in Kuala Lumpur is a good example of where what the software may see has to be agreed before the build, not during it.
Practical masking steps
- Replace names and phone numbers with consistent placeholders
- Remove identity numbers unless they are essential
- Keep test data separate from production data
- Agree how long shared data is kept, and how it is deleted
5. Who can confirm when an AI answer is right?
This question gets skipped most often, and it matters most. An AI feature is only useful if someone who knows the business can look at its output and say “yes, that is right” or “no, that is wrong, and here is why”.
That person is usually not the most senior. It is the merchandiser who knows how products should be described, the customer service lead who knows the correct answer to a refund question, or the supervisor who knows what a real defect looks like. In AI candidate screening, it is the hiring manager who checks whether the shortlist matches the role.
What we ask of them
- Review a set of real AI outputs each week during the build
- Mark each one right, wrong or unsure, with a short reason
- Help decide which mistakes are acceptable and which are not
An hour or two a week is usually enough. Without this person, nobody can say whether the AI is good enough to launch.
What happens next
Once you have answers to at least three questions, a project can start. The remaining questions are worked through in week one of a 4-week AI MVP, alongside user interviews and the first prototype.
If you are not sure how your data measures up, send us a short note describing what you have. We will tell you which questions are already answered and which need work.
Frequently asked questions
How much data do we need before starting an AI project?
It depends on the task. Many useful AI features, such as answering questions from documents or tagging products, work from what you already have. What matters more is that the data is reachable, reasonably consistent and checked by someone who knows the business.
Our data is spread across spreadsheets and WhatsApp. Can we still start?
Yes. Scattered data is normal for SMEs. The first step is listing where it lives and who owns each source, then choosing the one or two sources the first version actually needs.
Do we need to remove personal data before sharing it?
Share only what the AI genuinely needs. Where personal data is not needed, mask or remove it before it leaves your systems, and check your obligations under the data protection law where you operate.