6 May 2026

Is Your Power BI Data Ready for Copilot? Ask It Thirty Questions First

Copilot does not fix a semantic model, it inherits one. Thirty real business questions with verified answers turns AI readiness from an opinion into a pass rate you can watch move, and surfaces the capacity cost before the bill does.

"Should we turn Copilot on?"

I get asked this in nearly every room now, usually by someone who has already been told by someone else that the answer is obviously yes. And the honest answer is that it is the wrong question by one step. The question is whether your semantic model will give Copilot correct answers, because Copilot does not fix a model. It inherits one.

That distinction matters more than any feature comparison. A person opening a badly built report will squint at a figure, decide it looks off, and go and ask somebody. Copilot will not. It will answer in a confident sentence, in seconds, with no hedging and no trace of where the number came from, and somebody will paste that sentence into an email to a customer.

If the foundation is wrong, AI does not create a reporting problem. It broadcasts one.

The short version: test readiness rather than guess at it. Write thirty real business questions with your team, work out the correct answer to each one from source, then ask your model those thirty questions and count how many it gets right. That pass rate is your readiness score. Run it again after you fix things and you can see the number move. Everything else about AI readiness is opinion.

What Copilot is actually reading

People assume Copilot reads their data. It reads their model, and mostly it reads the parts of the model nobody bothers with.

Table and column names, as written, including the ones still called Sheet1 or tbl_fact_final_v2. Field descriptions, which are usually empty. Synonyms, so it knows that turnover, revenue and sales mean the same thing in your business, which are almost always empty. Measures and their names, so two measures called Revenue and Total Revenue that disagree by four percent become a coin toss. Relationships, including the ambiguous ones that give a plausible answer down one path and a different plausible answer down another. And hidden fields, or rather the ones somebody meant to hide and did not.

None of that is glamorous work. All of it is the difference between a useful answer and a confident wrong one. A model built for a human who already knows the business is not the same as a model built for something that knows nothing except what you wrote down.

Build the thirty questions

This is the part I would do even if you had no intention of switching Copilot on, because it is a decent test of a semantic model full stop.

Sit with the people who ask questions for a living, the finance team, the ops lead, the sales manager, and collect thirty questions they genuinely ask. Not clever questions designed to trip it up. The dull real ones. What did we sell last month. Which customers are over their credit limit. How does this quarter compare to the same quarter last year. Who are our ten biggest customers by margin rather than revenue. Which jobs went over on labour.

Then, and this is the bit people skip, work out the right answer to each one yourself, from source, and write it down. Thirty questions with thirty verified answers is a harness. Thirty questions without them is a demo.

Now ask the model. Score each answer as right, wrong, or right by luck, and pay attention to that third category, because an answer that happens to be correct while the reasoning underneath it is wrong is the one that will bite you in three months when the data shifts.

You will get a percentage. In my experience the first run is a lot lower than anybody in the room expects, and the failures cluster: dates, anything involving a ratio, and anything where two measures could plausibly have been the one meant.

Fix, re-run, and watch the number move

Once you have a pass rate, readiness stops being a matter of opinion and becomes a project with a metric on it. Rename the ambiguous fields. Write the descriptions. Add the synonyms your business actually uses. Delete or hide the four abandoned measures. Resolve the ambiguous relationship paths. Then run the same thirty questions again.

The re-run is the valuable half. It evidences the improvement rather than claiming it, and it gives you something to show whoever signed off the budget. It is also the natural thing to re-run after any significant model change, forever, which is a nice side effect of having built it once.

The cost conversation, which nobody has early enough

Two numbers, and the gap between them is where budgets get embarrassed.

Copilot in Power BI needs Fabric capacity, with an F2 as the floor. Sustained real use, by real people, all day, pushes well above that floor. It is very easy to pilot happily on something that will not carry production, sign off the pilot, and then discover the running cost at the point where you have already promised it to the business.

Model the capacity against your actual refresh patterns and user counts before the pilot, not after. Have the finance conversation before the bill does it for you.

What to do when the answer is that you are not ready

Then that is the finding, and it is worth considerably more than a green light would have been.

It is also not a no. It is a list, usually a short one, and most of the work on it is unglamorous model hygiene rather than anything architectural. Names, descriptions, synonyms, one measure of record for each figure, no ambiguous paths. Weeks, not quarters, on most estates.

The other output worth writing down is who should not have Copilot access yet, and why. Not as a punishment. Some parts of a model are ready and some are not, and giving natural language access to the half that is not ready is how you generate a confident wrong number in front of a customer.

Where to start

If you want to try this yourself, the thirty questions and the verified answers are the whole method, and you can do it with a spreadsheet and a couple of hours from each of three people. It is genuinely worth doing before you spend anything.

If you would rather have it run properly and scored, that is what the fixed-price AI-Ready Power BI Assessment is: models scored, the harness built with your team, capacity costed, and a ranked list of what to change first. And if the answer you actually need is whether Fabric is the right move at all, the Fabric Readiness Assessment is the one that answers that, including when the honest answer is not yet.

AI sits on top of the build, not instead of it. That is not a reason to wait. It is a reason to spend three weeks on the model first.

Stop arguing about the numbers. Start using them.

Book a fixed-price Power BI Health Check and find out how trusted, usable and audit-ready your reporting really is, and exactly what to do next.