How I Use AI as a Thought Partner
Concrete examples of my resource creation process
WARNING: This newsletter really is for my fellow super geeks. It’s long, technical, and about as niche as newsletters get. I not only go deep into my maths resource creation process using Claude, but also delve into the wonderful world of Atomisation and the sequencing of positive and negative examples. If you stick with it until the end and it finishes you off, don’t say I didn’t warn you…
Once upon a time
When ChatGPT first arrived, I couldn’t believe my eyes. You could name a topic — adding and subtracting fractions, say — and out came a worksheet that looked half-decent. I made a lot of those worksheets… and learned absolutely nothing in the process.
That is trap number one: cognitive outsourcing, to use Paul Kirscner’s phrase. The AI does the thinking, you copy and paste, and your own pedagogical reasoning slowly atrophies.
When I eventually realised what I was doing, I overcorrected.
My ego was the issue. Twenty years in the classroom and ten years of reading and writing about variation theory had convinced me that no AI was going to teach me anything about question sequencing. So I would design every worked example, every sequence, every variation myself, and use the AI only to format it nicely. Pretty layout, neat margins, well-formatted equations, done.
That is trap number two. It’s nowhere near as serious as the first because you are doing all the thinking. But it robs you of an opportunity to use AI’s potential to help you become a better planner and resource creator.
What I am doing now sits in between. I use AI as a thought partner.
But what the heck does that mean? Well, in this newsletter, I’ll try to explain, using the framework that I employ and concrete examples from my own resource creation over the last week.
My five-step framework
So how do you turn an AI from a solo worker or an order-taker into a thought partner? There are five steps I follow.
1. End every prompt with “Before you build, ask me some clarifying questions”
Tag this onto the end of any prompt, and the experience changes. You are signalling that you do not want passive execution. You want feedback.
So, the prompt becomes: “I want a worksheet on adding and subtracting fractions. Before you build, ask me some clarifying questions.”
The response will not be a worksheet. It will be a series of questions that clarify in the AI’s mind — and, more importantly, in your own mind — what exactly you want. The sub-skills covered, the difficulty, the progression, the sequencing, the style of questions.
This is the easiest move in the whole framework. You can do it on your next prompt without changing anything else.
2. Invite the AI to be a critical collaborator
Step 1 will get you good clarifying questions. Step 2 will get you a peer.
At the start of your next prompt, define the AI’s role and the rules of engagement:
“You are an expert on maths pedagogy. I am inviting you to a co-construction session to design a high-quality teaching resource for Year 8s to practice adding and subtracting fractions. I don’t want you to just accept every suggestion I make. I want you to challenge me with reasons when appropriate. Here is what I am thinking… [tell the AI what you are thinking]… Before you build, ask me some clarifying questions.”
The crucial line is the one in the middle: “I don’t want you to just accept every suggestion I make. I want you to challenge me with reasons when appropriate.”
Without that, the AI’s default is to agree with you. With it, you give the AI permission to behave like a peer — to push back when it sees a problem, to flag a misconception you are about to plant, to suggest an alternative where one exists. But crucially, not just to disagree for the sake of disagreeing (we’ve all worked with people who do that, right?).
3. Scrutinise the first output the way you would scrutinise a colleague’s
Once you have answered the AI’s clarifying questions and it has produced its first version, treat it like a draft a colleague has shared in the office. Pick out what works and why. Pick out what does not and why. Suggest specific alternatives.
Something like: “I liked your choice of denominator for question four — it reduces the arithmetic burden, so the student’s working memory can stay on the concept of adding fractions, not on the arithmetic. But your jump from question five to question six is too steep. A better question six would be [X]. What do you think?”
The “and why” matters — vague feedback gets vague rewrites. So, does the “What do you think?” — it once again reminds the AI of its role, not as a passive producer but as a collaborative peer.
4. Ask for a second version. Then a third. Then maybe a fourth.
The second version will be much better than the first. But it is rarely finished. Expect another round or two of feedback before the resource is where you want it. This is normal. This is also where the real cognitive work happens — for both of you.
When I do this, the final resource often ends up better than what I would have produced on my own. Partly because I am drawing on the AI’s expertise. But mostly because I am freed from the grunt work — the typing, the formatting, the layout — and that frees up working memory to notice things about the pedagogy I would have missed if I had been doing the whole job myself.
5. At the end, ask: “What did you learn from this collaboration?”
This is the one most people skip. After a long, back-and-forth, with a well-deserved, lovely resource at the end of it, the temptation is to close the chat and get a coffee. Don’t. Ask the AI what it learned that it could apply to future collaboration sessions. It will go back through the conversation, look at what you praised and pushed back on, and return with a set of principles.
Save those principles. Paste them into the start of the next chat. The next resource will be better, quicker.
To take this to the next level, apply the same trick to yourself. After a resource is finished, ask yourself what you learned. Update your own running document. The two documents, yours and the AI’s, should compound across resources.
Before the examples: one more thing on prompting
Before we get into the nitty-gritty, one more thing.
I no longer type my prompts and responses. Typing on a keyboard is so March 2026.
I now dictate my prompts and responses using a tool called Wispr Flow, which transcribes voice into any text field on my computer. I mentioned this on my recent podcast with Becky Allen, and as a result of people clicking the link, I’ve got more free months of Wispr Flow use than science would suggest I’ll be able to use in my own lifetime.
There are two reasons Wispr Flow has changed how I work.
One, it is quicker than typing. Much quicker.
Two — and this is the bigger one — it frees up attention. When I am typing, a chunk of my working memory is on the keyboard. When I am dictating, all of it is on the idea. I can chase a half-formed thought before it slips away. I can include the kind of throwaway context — “what worries me here is…”, “and I think the reason is…” — that I would strip out if I were writing. The more I read about how large language models work, the more I realise they are hungry for that kind of context. The flowery asides we filter out when typing turn out to be exactly what sharpens the output.
So whenever you read one of my quoted prompts below, picture me talking into my computer, while my wife once again sighs, knowing that I am chatting to my best mate, Claude.
Five examples from last week
Strap yourselves in, we are about to get super nerdy. You have been warned
This week, I’ve moved on from my Shedloads of practice resources (in case it is useful, my last few builds were surface area and volume, area and circumference of a circle, representing linear inequalities on a graph, and an epic ratio page, along with this circle theorems interactive tool) to experimenting with designing teaching and testing sequences for categorical atoms.
The idea comes from my good friend, Kris Boulton, over at the excellent Unstoppable Learning Substack, which itself builds on the pioneering work of Siegfried Engelmann on direct instruction and logically faultless communication.
I will spare you the full theory (Kris is the person to read on this, starting with this post). The short version: when you want to communicate a categorical concept precisely — what counts as an integer, what counts as an exterior angle, what counts as a valid set of angles on a straight line — a carefully chosen sequence of non-examples and positive examples lets students induce and encode the rule for themselves, without you ever stating it.
I have been writing these sequences by hand for three or four years. I find them intellectually stimulating. I also find them hard. So when AI’s diagram capability finally cracked (I wrote about this here and here), I wondered: could AI help me build the animations needed to draw students’ attention to the minimal changes between positive and negative examples, and help me sequence these positive and negative examples in the first place?
What follows are five exchanges from recent builds. The first three are from an “Is this an integer?” sequence. The fourth is from an “Are these angles on a straight line?” sequence, and the last from “Is this written in standard form?”. All five exchanges shaped the final resource in ways I would not have got to on my own… and that the AI would not have got to on its own.
Exchange 1: How many examples do we need? (the AI pushes back)
I had drafted a sequence of five examples to try to convey to students as clearly as I could: “What is an integer?” My prompt to Claude was:
Perhaps we start off with a decimal like 0.2 for our first non-examples, and then show that transitioning to an integer, 2.
For our second positive example, we show a really big integer, to try and push the boundaries of it, so something like 53,076, or something crazy like that.
For the third positive example, we need to show them a negative number, so maybe something like -3, and then we could transition that into -1/3 for the final non-example.
Would that cover everything, or do we need to also show them that a negative decimal is not an integer? I think the five examples might be enough on their own. Let me know what you think.
What I expected was: “Sounds good, Craig, you are a genius. Let me draft the sequence.” What I got was:
I’d push for adding -2.5 as Example 6.
After Example 5 the student has seen one positive decimal as N (0.2) and one negative fraction as N (-1/3). To classify -2.5 as not-an-integer they need to combine two threads from the sequence — ‘decimal means not whole’ and ‘negative is OK only when whole’ — and apply both to a single new case. That’s the kind of inference that works for the strong half of the class and fails for the rest.
A KS3 student who has internalised ‘fractions like -1/3 aren’t integers’ can still trip on -2.5 because decimals feel categorically different from fractions at that age, even though we know they’re the same thing.”
I had to stop and think.
The argument was right. I had been reasoning from my own perspective — for a teacher who knows that -2.5 is just another not-an-integer, of course, five examples is enough. But the sequence was not for me. It was for a Year 8 student in the bottom set, who may have noticed that decimals are not integers, and negative fractions are not integers, but negative whole numbers are. They have not yet been given any signal as to which category negative decimals belong to.
We went with six examples.
The bit that mattered was not the conclusion. It was the pedagogical reasoning that the AI shared. Claude did not just disagree — it surfaced something that my curse of knowledge blinded me to.
Don’t say I didn’t warn you: this post is for the super geeks.
Exchange 2: Two ideas I really wanted to use (the AI pushed back again)
Once we had agreed on six examples, I got excited about the closing transition.
My first idea was to use -1.3 as the sixth example, transitioning from -1/3 in the fifth. The reason: a really slick animation. With -1/3 written as a stacked fraction, you could fade out the dividing line, slide the denominator down, and morph it into -1.3. Minimal change. Visually clean. I was pleased with myself.
Claude was not.
I’d reject this on a stronger ground than ‘no mathematical relationship.’ A student who watches -1/3 morph visually into -1.3 — bar vanishes, denominator slides up to become a decimal digit — could plausibly encode ‘so -1/3 must equal -1.3’ as a takeaway. That’s an active misconception we’d be planting in service of a slick animation. Bad trade.
Claude doesn’t beat around the bush.
But Claude was correct. I had been so focused on the visual elegance that I had missed what the student watching it would actually take away.
So I tried again. My second idea was to choose a negative fraction and pair it with its decimal equivalent — same value, two surface forms, both not-integers. Pedagogically tidy. Again, pushback:
I’d skip this. The pedagogical payoff is ‘watch one number in two surface forms, both not-integer,’ but the student has to compute the equivalence to receive the payoff — recognise that -3/2 = -1.5. If they don’t, they see two different examples and the equivalence does no work. For an atom whose entire job is integer/not-integer, asking the student to do parallel fraction-to-decimal arithmetic distracts from the main question.
Two ideas, both rejected. Both, on reflection, rightly.
I had been thinking about what looked good and what was conceptually clever. The AI had been thinking about what the student would actually perceive and what cognitive load would be wasted. The student-side reasoning was the right reasoning all along — and I had missed it twice in a row.
Exchange 3: My turn to push back (the negative sign needs to physically stay put)
The first two exchanges could make the AI sound like the wise sage. So here is one going the other way.
We had settled on the closing block of the integer teaching sequence:
I needed students to notice that the negative sign had three different things after it, resulting in two different answers. The minus is the constant. What follows the negative sign is what does the work.
But the way Claude had set up the animations, the minus sign was redrawn every time the example changed and moved based on the length of the number that followed it — not held in place as a fixed anchor the student could rely on. This reduced the likelihood that students would notice it.
This might look like a design quibble. It was not.
Claude, I’m not happy with the transitions through the closing block. The whole pedagogical point of E4, E5, E6 is that the negative sign is the invariant — same minus, different operand. But if the minus sign on the screen disappears every time the example changes and gets redrawn, the student cannot see that. The negative sign has to stay in the exact same position throughout. Visibly. The whole time. The thing-after-it can morph or fade, but the minus does not move. And I think this is a general principle we need to bake into the handoff document: any constant feature that survives a minimal change must remain in its exact same position and visible at all times.
The reply was a clean concession, with a useful refinement:
You’re right and the argument is strong. The whole point of a minimal-change transition is that the student can read off what changed by reference to what stayed the same. If everything in the candidate visually leaves and gets replaced — even by an identical character — the student loses the anchor that lets the change land as minimal. One refinement: pieces should be identified by pedagogical role, not by character coincidence. The ‘3’ in ‘-3’ and the ‘3’ in ‘-1/3’ share the character but not the role — operand body in one, fraction denominator in the other. Linking them would require motion to morph one into the other, which isn’t invariance — it’s an additional animation. The minus is the right invariant. The bodies of the operands should cross-fade as wholes.
That refinement mattered. I had been about to ask whether the ‘3’ should be linked across examples too? But Claude was right. The ‘3’ in ‘-3’ and the ‘3’ in ‘-1/3’ are different mathematical objects, even if they share a glyph. Treating them as the same piece would have risked planting a quieter misconception.
What might look like a fiddly design decision was, in fact, the load-bearing pedagogical move of the whole closing block. The negative sign, held still on screen, tells the student: this is the rule. Without it, the three examples could read as unrelated, with reasons the student has to guess. With it, they read as one rule applied to three operands.
Would I have reached this conclusion on my own? Maybe. But perhaps my attention and working memory reserves would have been depleted with all the other aspects of the sequence I had to contend with. Whereas with Claude as a thought partner, some of those reserves were freed up, enabling me to spot things like this
Exchange 4: How much do we need to teach?
So far, three exchanges from the same number-focused build. As regular readers will know, I’m currently obsessed with making diagram-heavy resources using AI. So, here is one from a geometry resource: “Are these angles on a straight line?”
The critical feature for this sequence is that the marked angles must tile exactly half a plane on one side of the line. Two of the ways this can go wrong are:
Overfill: the student includes an extra angle that spills past 180°
Gap: the student marks angles that fall short of 180°, leaving a gap before the line
Both are failure modes of the same rule. Both produce incorrect answers for similar reasons. So the question was: do we include both in the example sequence — two closing non-examples, one of each — or do we teach one and trust students to generalise to the other?
Claude proposed the latter. Include one of them in the teaching sequence (the one students are less likely to spot on their own), and give both equal weight in the testing sequence that follows.
The reasoning was about scope. The teaching sequence is finite. Every example has a cost. Some failure modes need to be shown explicitly because students will not get to them on their own (like the negative decimal from the integer sequence above); others will land once the rule is understood, and adding an example for them just bloats the sequence without paying off. A critical feature does not require a teaching example for each failure mode. It needs a teaching example for the failure mode that students will not spontaneously generalise to.
I would not have arrived at this distinction on my own. I would have over-engineered — teach everything explicitly, leave nothing to chance. The AI’s framing — teach what they cannot generalise, test what they can — is sharper. And the resulting sequence is shorter and better.
The principle might still be wrong, of course. Maybe students do not generalise to overfill as easily as we think. Likewise, in the “Is this an integer sequence?”, maybe I could have got away with just five examples in the teaching sequence. The testing sequence will tell us. If gap performance is much worse than overfill performance after teaching, we add a gap example to the next iteration of the teaching sequence. That is the kind of judgment a thought partner helps you make. You then gather data from real students in real classrooms and adapt accordingly.
Exchange 5: When a non-example fails too hard
Anyone still with me?
Okay, one more example, this one is hot off the press, from a “Is this written in standard form?” sequence that I wrote last night.
Here is the teaching sequence Claude initially proposed:
Notice the final example. This is not in standard form for two reasons: the coefficient is out of range, and the exponent wasn’t an integer.
The intention was efficiency: get a final example in on both critical features at the close. The pedagogical effect, I realised, was the opposite.
We have a big problem with example six. We’ve already communicated to the students that standard form fails if the power is not an integer. Students may rightly look at example six and think, “I don’t even need to pay attention to that starting number because the power is not an integer, so it must not be standard form.”
Claude got it immediately:
The double-failure in E6 doesn’t just risk a “two negatives” rule-induction — it actively lets students off from processing the coefficient. They see the non-integer exponent and stop reading. The headline critical feature gets bypassed at exactly the slot where it should land hardest.
The deeper point that generalises is the difference between “this example fails on one thing” and “this example presents one new thing for the student to process.” These are not always the same. If you’ve already shown a student that a non-integer exponent disqualifies an expression, and you then show them another expression where the exponent isn’t an integer and the coefficient is also out of range, you’ve created a situation where the student can land on the right verdict without ever reading the coefficient. The critical feature you wanted to teach gets skipped, even though it’s there on the screen.
That’s the kind of thing I need to be thinking about when resource creating - and the very thing I might have missed if I had been bogged down with the minutia of creating the rest of the sequence.
At the end of each of these five exchanges, I asked Claude to reflect upon what he had learned and then write up the principles in a document that I could share at the start of every subsequent page build.
The resources
This is very much a work in progress. I have not added them to my Topics page yet. But in case you want to try out the resources with your students, here are the ones that I (well, me and Claude) have built so far, complete with animations:
One last thing: Is this all a bit sad?
You might be reading this and thinking: Come on, Craig. This is just collaborating with a colleague… except the colleague is a computer, you weirdo.
My wife would agree.
And, look, I would not pretend that bouncing ideas off Claude is better than bouncing them off a great colleague with a mug of tea in the staffroom at the end of a long day. Of course it is not. A human colleague may have years of context on your students, your school, your department, and on you as a teacher. A human colleague can read the room. A human colleague gets the joke.
But here is the key question: How many of us, in our particular school, in our particular department, in our particular subject, have a colleague genuinely up for this kind of conversation — a back-and-forth on the pedagogical structure of a six-example sequence, with multiple back and forths, and who will do most of the grunt work? Oh, and who is available on a Sunday evening, or when the footy is on?
I am very lucky to know quite a few of those people. Most I met through writing and the podcast, including Kris Boulton. But even Kris is not available 24/7 (23/7, maybe).
Claude is.
That is not better than the staffroom conversation. But it is much better than not having the conversation at all.
Over to you
You made it to the end! You deserve a medal.
But instead, I have a challenge: If you have given AI a go as a thought partner — actually asked it “Before you build, ask me some clarifying questions”, actually asked it “What did you learn from this collaboration?” — drop your examples in the comments. The dynamic I have described above is the thing I most want to hear about. Not the prompts that worked the first time, but the pushbacks that made you stop, think, and learn something.
My wife thinks I have lost it. She might be right. But I have never had this much fun making maths resources.
Thanks so much for reading, and have a great week.
Craig
PS You can push AI too far. Here’s how Claude responded to one of my recent suggestions after an hour of back-and-forths:
AIs are getting increasingly human by the day.













This is my first Substack comment. I was really excited reading this article and actually got through to the end! What struck me most is the way you have crafted AI as a thinking partner that helps you by encouraging it to push back constructively. That's blowing my mind.
I do lots of one-on-one tuition, and the way I use AI at the moment, or the way I found it to be the most useful, is: after a lesson's finished I will run my impressions and thoughts of the lesson through it and use it as a thinking partner to help me track misconceptions and see where deep conceptual issues are surfacing as minor mistakes and reasoning gaps. I've been finding it recently very helpful, but also I've been getting worried about its excessive praise. So now I'm ready to try a more honest and combative AI model and I'm really excited about it! Thank you so much for sharing this.
Thanks Craig, really useful discussion. I think it hits at the core of successful working with AI and that is collaboratively. AI is now demonstrating when used effectively the capacity to be the critical friend that social media should, but often fails to be. For the less confident teacher to be able to discuss pedagogy without judgement is a really positive step forward.
Please don't frame this as nerdy, we need these conversations to be normalised in the same way we resist the phrases "this is hard and I wasn't good at maths" as acceptable default positions.
When we get excitement at writing the questions, pupils are more likely to be excited by the answers. Perhaps that needs to be a prompt...create an question with a surprising answer! I'll investigate.