20 Comments
User's avatar
Jack Morrice's avatar

This is my first Substack comment. I was really excited reading this article and actually got through to the end! What struck me most is the way you have crafted AI as a thinking partner that helps you by encouraging it to push back constructively. That's blowing my mind.

I do lots of one-on-one tuition, and the way I use AI at the moment, or the way I found it to be the most useful, is: after a lesson's finished I will run my impressions and thoughts of the lesson through it and use it as a thinking partner to help me track misconceptions and see where deep conceptual issues are surfacing as minor mistakes and reasoning gaps. I've been finding it recently very helpful, but also I've been getting worried about its excessive praise. So now I'm ready to try a more honest and combative AI model and I'm really excited about it! Thank you so much for sharing this.

Craig Barton's avatar

Thanks Jack, that's a great use case. Setting the AI up to not always agree with you really does make a difference, but you do need a thick skin because it doesn't hold back!

Jack Morrice's avatar

oh my gosh, I've just made the change and i see what you mean!

Andy Parkinson's avatar

Thanks Craig, really useful discussion. I think it hits at the core of successful working with AI and that is collaboratively. AI is now demonstrating when used effectively the capacity to be the critical friend that social media should, but often fails to be. For the less confident teacher to be able to discuss pedagogy without judgement is a really positive step forward.

Please don't frame this as nerdy, we need these conversations to be normalised in the same way we resist the phrases "this is hard and I wasn't good at maths" as acceptable default positions.

When we get excitement at writing the questions, pupils are more likely to be excited by the answers. Perhaps that needs to be a prompt...create an question with a surprising answer! I'll investigate.

Craig Barton's avatar

Great point about the framing, Andy. And the analogy to "I wasn't good at maths". I always try and celebrate my nerdiness and geekiness, but I'll be much more explicit next time. Let me know how you get on with your experimentation about surprising answers.

James Cantonwine's avatar

This is brilliant, Craig. I really appreciate seeing the details of how other people work with AI, especially Claude. It really helps sharpen my own thinking about the process I use.

I really should give voice interactions another try, but I find that typing helps me to avoid cognitive outsourcing. It pushes me to get my own thoughts out in detail, review them, and then send them on. I'm a little leery of going all in on dictation, rambling for a bit, and then just accepting whatever came out.

At this point, I've explicitly asked Claude for pushback so many times that this is the last sentence Claude has in its managing memory for me: "His thinking is characterized by a willingness to engage with and push back on intellectual challenges rather than seek validation." While I no longer need to invite Claude to have it push back, I've found it helpful to think about what bias I'm trying to mitigate and clearly communicate that. Generally it's confirmation bias - I've had some idea that I think is just wonderful and am looking for Claude to tell me where it's going to fail so that we can build around that.

As for collaborating with colleagues, I think Ethan Mollick's Best Available Human essay from 2023 is still top notch: https://www.oneusefulthing.org/p/the-best-available-human-standard. AI doesn't need to be as good as collaborating with the best colleagues, it just has to be better than whomever would actually be available to you at that moment. And it's definitely better than me sitting at a desk by myself!

Clive Long's avatar

I'm experimenting with AI to rapidly build apps to tackle A-Level topics or questions that students find challenging based on class questions. I'm using Claude and am going to try Codex. I have a 10 dollar a month subscription for both. I regularly run into usage limits for Claude - but that's ok because I tend to work in bursts. the code is all held on Github and served using Git Pages so the apps are limited to using HTML/CSS/JS and sometimes React. A bit of a learning curve to understand this tech but I also teach CS. Some examples of what Claude has produced are :

https://maths-apps.crlong.uk/Stats/nested-distributions.html

https://maths-apps.crlong.uk/Stats/practise-probs/nested-distrib-test.html

https://maths-apps.crlong.uk/Mechanics/practise-probs/@mech-practise-home.html

An advantage of this setup is that when I find errors or think up "improvements" to the app , I just download the code using Github desktop, tweak it by hand or using Claude AI, test locally, then reupload to GitHUB and the site is refreshed in a few seconds with no outages or changes to page links. Next step is to try to measure how effective this stuff is. I don't want to reproduce the hundreds , thousands of sites covering the same topics. I think the next step is to make sure the presentation is optimized for mobile device. If it looks bad on a phone it won't be used.

RachelCT's avatar

Really interesting read and I’m not even teaching maths at the moment. I’ve been tinkering with using Claude and Gemini to generate a first draft of sets of questions for computer science practice before editing them myself but this has given me ideas for a more useful process of collaborative refinement.

Craig Barton's avatar

That's great, Rachel. Good luck with it. Let me know how you get on.

Tony Purkiss's avatar

Well, as a super-nerdy geek I loved this post! I have been part of a working group this year looking at AI and building resources and your series of posts have been really helpful to me. I have not used Claude as yet, but it seems for maths that it's one (or many?) step ahead of its rivals at the moment. Copilot and ChatGPT have improved immeasurably since last September but I cannot get them to reliably produce resources like I have seen you achieve in your posts, but this is definitely down to my novice, but improving, approaches. Getting those prompts right is the one big lesson I have learnt.

Craig Barton's avatar

Yes, I keep going back to Gemini and ChatGPT, but I can never get the outputs as good as I can for Claude. Also, I just enjoy the interaction with it. The inescapable fact, though, is that I think you probably need a premium plan in order to get these high-quality outputs. The free models are certainly getting better, but they just don't cut it. Then you run into an even bigger problem: as soon as you pay for the premium and you get right into the middle of a resource-creating session, you hit your limit. The dilemma is, can you wait three hours for the limits to reset, or do you upgrade again??????

Tony Purkiss's avatar

I have upgraded with Copilot but still don't think it can match what I see from the outputs you get from Claude, but again that may be down to my prompts/feedback etc? Can I wait 3 hours...I can but it is frustrating when you are almost there with exactly what you want! If there was no limit then I think that my wife might conclude that I've been kidnapped as I disappear down the AI wormhole...

Gail Brown's avatar

Hey Craig, I've been following you and reading your Friday posts in particular - and am SO GLAD that I've just read this post! And, I also have Paul Kirschner as one of my long term heroes!

This is exactly the type of idea I have - except for your Math content... I want to do this for reading and for learning... and for younger students, so some of your language is, I think, more for secondary students along with your Math content? Please correct me if I'm mistaken here...

And Math isn't really my strength - so some of your logic for the specific examples did go "over my head! I agree that AI might "match"/"search" for solutions better - I think AI has UNLIMITED working memory - unlike us Humans!

I like most of your steps, and I'd challenge you back, with your own conclusions:

1. AI is a computer - so - I don't believe it "thinks" - so I'd change some of your wording to reflect that AI is a searching / mirroring tool? For example, and probably in agreement with your wife (??) I would never use a personal pronoun with AI - to confirm it is NOT a human?!!

2. You make the point about AI not being what I would call "a real colleague" - which I agree with! So, why couldn't you build in an optional step in your strategy sequence where you a "take time out" (to think?) and maybe find a REAL HUMAN colleague and ask them? NOT sure if this is an option for your examples, though I think it might work for some of my literacy example...

3. I believe in teaching for generalisation - so I think we agree again - so I think I'll try your steps with some of my specific examples for literacy, and see if AI does come up with "better solutions"...

My work includes an "overarching conceptual framework" which provides a foundation for the limits of generalisation??

I'm NOT sure whether I should include this framework - or leave it out and see what AI does?

4. Lastly, you've used CLAUDE - do you think you'd get the same sorts of results with Gemini???

I will try your strategy steps with some of my content - and compare Claude and Gemini as I do that, with the same prompts?

And I've be really keen to hear your feedback on my suggestions, and whether any of these options might be worthwhile??? Please let me know? Many thanks, again! Gail :)

Craig Barton's avatar

Hi Gail. Thanks for your comments. I agree with all of them:

1. I find it hard to avoid saying things because I like reading the process that CLAUDE goes through before the output comes, and I don't know another verb to describe what I'm seeing.

2. I like point two about finding the real human, so long as you're privileged enough to be able to find one who's willing to talk about the thing you want to talk about at the time you want to talk about it!

3. When it comes to CLAUDE, I've tried Gemini and I've tried ChatGPT, and for the time being I can get nowhere near this level of interaction and quality of output.

Thanks again for your comments.

Gail Brown's avatar

Agreed - Point 2 - might work for students in a classroom - they might find another student - and ask or just compare their results for similarities & differences… or ask their teacher…

I’m going to put something together - and I think you (& others on Substack!!) are modelling what teachers should do AND what I would want to see as a teacher: Evidence of student learning WITH repeated prompts & examples they refine!

MANY Thanks again!! Ps I’m based north of Sydney Australia 🇦🇺 so my posts and feedback probably don’t fit your Timezone at all?! 🤣❤️❤️

Susan Knopfelmacher's avatar

I’m not a maths teacher, but really enjoyed reading this. The interaction is really telling. Clearly, you prefer Claude as a platform - for particular reasons?

Craig Barton's avatar

I’ve tried them all, and (for now) I find both the collaboration and the eventual output far superior. Thanks for your comment! At least I know 1 person read it!

Susan Knopfelmacher's avatar

Thanks for replying. All your posts are great, I’ve been following the AI Interviews with interest. My area is language/humanities but having a broad view is really useful. Again, thanks for all your work :)

Clive Robert LONG's avatar

Maybe not the exact entry for this comment ... but as in another post of mine I'm experimenting with Claude AI to produce apps that explain, animate and check understanding of topics that A-Level students have difficulty with. Sometimes I'm quite prescriptive in my prompts, sometimes I let Claude propose the topics, content and build the app. I was checking with a keen student one calculation as part of "nested probability", https://maths-apps.crlong.uk/Stats/nested-distributions.html, and she found an error in the calculation of a binomial probability. The lead up and explanation generated by Claude was I thought clear and well-paced.

The "lead up" material was (in summary)

Y∼B(12,0.6284)

P(Y=9)

Yielding

=220×(0.6284)^9 ×(0.3716)^3

all fine

When this expression put in calculator, or using Binomial PD directly, this should result in 0.17249 ....

But the app displayed: P(Y=9)≈0.189

How this came about, I don't know, but looking at the not too complicated HTML the "answer" was statically coded

<div class="answer">\(P(Y = 9) \approx 0.189\)</div>

Now I'm trying to get the apps to randomly generate values and parameters and calculate the answer in "real time".

This is probably obvious but if using AI to generate teaching or testing content, don't trust it.

Separately but similarly I input some handwritten scripts and the mark scheme. ChatGPT did an amazing job of interpreting some of the scrawls and produced some useful, relevant and detailed feedback. In other cases ChatGPT completely invented feedback for about a quarter of the questions where the student had written nothing (the famous hallucination).

So has current AI improved my productivity, I don't think so, but I'm concerned that as skills and time made available to check all its output is going to be reduced we are all going to be guided by electronic snake-oil peddlers.

Kathryn Naylor's avatar

Really interesting. But how did you go the step further and animate and set questions ?