Hands-On Tool Exploration
Preethi Chandrasekaran described the first time she used an AI writing assistant as "like having a really fast intern who has read everything but understood nothing about my industry." The email it drafted was fluent and well-structured. It also opened with a pleasantry she would never use, cited a turnaround time that was wrong, and called the recipient "valued client", a phrase her organisation had stopped using two years earlier. She fixed it in three minutes and sent it. "Still faster than drafting from scratch," she told me. "But I have to know what to check." That productive scepticism, built through direct experience, is the entire point of this chapter.
You cannot build that scepticism by reading about AI. You build it by using AI, observing carefully, and noting what surprises you. This chapter is therefore unlike the ones before it: it is not information to absorb, it is exercises to complete. Each one gives you a prompt to copy and run in your own tool, a variation to try afterwards, and a set of things to watch for while you read the output. What you learn comes from the difference between the two runs, and from the details you would have skimmed past if nobody had told you to look.
How to Use These Exercises
The workflow is the same every time. Read the exercise description so you know what you are testing. Open whichever AI assistant you have access to, whether that is a general-purpose chat tool, an assistant built into your productivity suite, or whatever your organisation has approved. Copy the prompt exactly as written and paste it in. Review the response and ask how good it really is, where it is strong, and where it is weak. Write down what you noticed. Then run the suggested variation or follow-up, and compare the two results: what changed, and why do you think it changed?
You will need about forty to fifty minutes and somewhere to take notes, whether a notebook, a document, or anything else that suits you. The notes matter more than they sound like they do, because you will forget the specifics within an hour, and the written observations become your personal reference for what this particular tool does well and where it needs watching. Do not rush. The value is not in finishing quickly; it is in noticing details and building intuition, which only happens at the speed of actually reading the output.
One practical rule: do not skip an exercise because you think you know what will happen. The surprises are the learning. A result that confirms your expectation tells you something. A result that defies it tells you more.
Exercise 1: Professional Email Drafting
This is the task most professionals reach for first, and it reveals a great deal about how context shapes quality. The objective is to find out how well your tool handles professional communication and how far you can move its tone. Start with a deliberately vague request: "Write a professional email declining a meeting request." Read the output. It will probably be reasonable but generic, with no names, no specific reason, and a tone that is formal but impersonal. That is your baseline, and it is worth keeping in front of you when you run the next version.
Now supply the situation: "Write a professional email from me (Preethi, procurement manager at a mid-size manufacturing firm) to a vendor named Rajesh at SupplyLink Technologies, declining his request for a thirty-minute product demo this month. We are in a budget freeze and cannot consider new vendors until Q2. Keep the tone warm but firm. Under 150 words." Note what changes: the specificity of names, the accuracy of the reason, the control over tone, and whether it obeyed the word limit. Ask yourself which version was closer to something you could actually send, and what you still had to change by hand.
Then try a harder emotional register, because declining a demo is low-stakes and apologising is not: "I need to send an email to a client who has been waiting a long time for their project to be completed. The delay was due to unexpected technical challenges on our end. I want to apologize, explain briefly what happened, and give them a realistic timeline for completion. The tone should be professional but warm. Write the email for me." Read the result against a short checklist. Is it apologetic without being obsequious? Does it explain the problem clearly without sliding into excuses? Does it contain all three of the pieces you asked for, meaning acknowledgment, explanation, and timeline? Would you send it, or does it need editing first?
Finally, paste the AI's own response back into the chat and send this follow-up: "Can you rewrite it to be more casual and less corporate? I want it to feel like it is coming from a real person, not a PR department." Compare the two. What did the model actually change, what did it leave alone, and is the casual version better? The pattern you are documenting across all four runs is straightforward: AI-generated professional communication needs context to be good, generic prompts produce generic output, and tone is far more steerable than most first-time users expect.
Exercise 2: Document Summarisation
The objective here is to test how well the tool extracts key information and condenses length without distorting meaning. Find a document you have been meaning to read, ideally two to five pages long: a report, a policy document, a lengthy email thread. If you would rather work from a shared text so you can compare notes with a colleague, use the sample paragraph below. Either way, paste the full text into the tool and ask for "Summarise this document in five bullet points, each under 20 words. Focus on what I would need to act on or decide."
Sample document. Renewable energy adoption has accelerated dramatically over the past decade. Solar photovoltaic costs have dropped by 90% since 2010, making solar competitive with fossil fuels in most markets. Wind energy capacity has more than doubled globally. Electric vehicle sales have grown exponentially, particularly in Europe and China. Battery technology has improved substantially, with new solid-state batteries offering higher energy density and faster charging times. Grid modernization efforts are underway in most developed nations, creating smart grids that can balance intermittent renewable sources more effectively. However, significant challenges remain. Energy storage at scale remains expensive and technically difficult. Grid infrastructure in many regions is outdated and insufficient for distributed generation. Supply chains for critical minerals (lithium, cobalt, rare earths) are concentrated in a few countries, creating geopolitical risks. Public acceptance of renewable infrastructure, particularly wind farms, remains mixed in some regions. Transmission infrastructure from renewable-rich regions to population centers is often inadequate. Despite these challenges, projections suggest renewables will comprise 50% of global electricity generation by 2050, up from 30% today.
Read the summary against the original. Did it capture the main points, and did it miss anything important? Did it include something that felt irrelevant? Was the balance right, so that a document covering both progress and obstacles came back as both rather than as one? Could a non-expert follow the summary on its own? And the check that matters most: is there anything in the summary that is not in the original? Invented detail in a summarisation task is an outright error, not a stylistic quibble.
Now run two variations. First change the format: "Now create a bullet-point summary instead, five key points I should remember from this document." Ask yourself which format is genuinely more useful to you and why. Then change the framing instead of the format: "I am a finance manager. What in this document is relevant to a financial decision or financial risk?" The second version will pull different content out of the same text. That is the lesson. The AI has no fixed idea of what "important" means; it responds to your framing. This is a feature, since your framing is what surfaces the relevant material, but it also means an unexamined framing produces a filtered summary without you ever realising it was filtered.
Exercise 3: Data Analysis and Pattern Recognition
Here you are testing whether the tool can read structured data, identify patterns in it, and be honest about the limits of what the data supports. Use your own numbers if you have a small table to hand, such as sales figures across regions, attendance across months, or survey responses. Otherwise use the sample below, which is ten months of sales data for a small business. Paste the table in as plain text, keeping the column headings, and give the tool a task with several distinct parts so you can see which parts it handles well.
| Month | Product A Sales | Product B Sales | Product C Sales | Marketing Spend |
|---|---|---|---|---|
| Jan | 45000 | 32000 | 18000 | 5000 |
| Feb | 48000 | 35000 | 20000 | 6000 |
| Mar | 52000 | 38000 | 25000 | 7000 |
| Apr | 51000 | 40000 | 28000 | 8000 |
| May | 54000 | 42000 | 30000 | 8000 |
| Jun | 56000 | 45000 | 32000 | 9000 |
| Jul | 55000 | 44000 | 31000 | 7500 |
| Aug | 58000 | 47000 | 35000 | 9500 |
| Sep | 60000 | 49000 | 37000 | 10000 |
| Oct | 62000 | 51000 | 39000 | 10000 |
Ask: "Here is 10 months of sales data for our company. Please analyze the trends and tell me: (1) Which products are growing fastest? (2) Is there a relationship between marketing spend and sales growth? (3) What patterns do you see? (4) What recommendations would you make going forward?" Then read the answer as a marker rather than as a reader. Did it correctly see that all three products are growing? Did it identify that Product C has the highest growth rate, which is the answer the raw totals hide? Did it notice that marketing spend moves with sales? Are the recommendations grounded in these numbers, or are they generic advice that would fit any dataset? Did it flag its own limitations, such as the small sample? And are there patterns you can see that it missed?
Two follow-ups make the exercise worth the time. The first tests extrapolation: "What would happen to sales if we doubled marketing spend? Can you project sales for November and December based on current trends?" Judge how confident you are in what comes back. Are these reasonable extensions of the data, or speculation dressed in the same calm tone as the analysis? The second follow-up is the more revealing of the two: "What data would you need to confirm whether your explanations are correct?" AI tools are generally good at spotting patterns and generating plausible explanations for them. They are much weaker at knowing what they do not know, and asking directly what they cannot confirm surfaces the gaps the first response treated as settled.
Note for your log whether the output distinguished between patterns in the data and explanations for those patterns. Strong outputs separate observation from inference and say which is which. Weaker outputs blend them, stating a guess about causes with exactly the same confidence as a number you can verify yourself in the table.
Exercise 4: Creative Brainstorming
Brainstorming is one of the clearest strengths of current AI tools, so this exercise is less about whether it works and more about how you should use what comes back. Pick a real problem you have been thinking about, something genuine: a process that frustrates you, a communication challenge, an initiative that has stalled. Run "I am trying to [describe the problem in one sentence]. Give me ten possible approaches, including at least two that are unconventional." Then count. How many of the ten had you already thought of? How many were genuinely new? Were any surprising enough to be worth exploring properly?
If you want a second data point with a problem you have no emotional investment in, try a scenario from outside your own work: "I run a local coffee shop in a medium-sized city. I want to increase foot traffic and build community. What are 10 creative ideas I could try? They should be realistic (not requiring huge budgets) but memorable. I want to attract both regular customers and new people." With unfamiliar territory you can judge the output more coldly. Are the ideas genuinely creative or a list of the obvious? Are they realistic for a small business or fantasy? Do any of them show that the model actually absorbed the constraints you gave it about budget and audience, and does it seem to know which of its own ideas are the better ones?
Then push one idea into depth, either with "Take idea number 4. What would be the three biggest obstacles to implementing it, and how might each be addressed?" or with "Let's develop [IDEA] further. What would be the first three steps to implement it? What could go wrong? How would I measure if it worked?" Watch whether the tool helps you think the idea through or stays at the surface while you do the real work. The value of brainstorming is rarely in using an idea wholesale; it is in using the list as a trigger for your own thinking. The idea you would never have considered is often the one that unsticks a problem you have been circling for weeks.
Exercise 5: Research and Information Synthesis
The final exercise tests research and synthesis, which is where confident-sounding output is most dangerous. Start with a factual question about your own industry or function, chosen so that you already know roughly what the answer should be and can therefore grade the response. Something in the shape of "What are the main regulatory requirements for data retention in [your country or sector]?" works well, because it is specific, it is jurisdictional, and you will notice immediately if the answer drifts into generalities or into another country's rules.
Then run a comparison question in an area where you are a genuine beginner, since that is the condition most people are actually in when they use AI for research: "I need to understand the key differences between React and Vue.js for a project decision. I have not used either before. Can you explain the main differences in a way a beginner can understand? Include: (1) learning curve, (2) community and ecosystem, (3) performance, (4) job market demand, (5) use cases where each excels. Keep it concise, I do not need deep technical details." Assess it on five things: whether it explained technical differences in accessible language, whether the comparison was balanced or quietly favoured one option, whether the information felt current or dated, whether it committed to a recommendation or hedged until it was useless, and whether you ended up better placed to decide or merely better informed about your own confusion.
Follow up in the direction a real decision would take you: "I have already decided to go with React. What are the best resources to learn it quickly if I am an experienced JavaScript developer? Books, courses, tutorials, what do you recommend?" Recommendations are where knowledge cutoffs bite hardest, so watch whether the suggestions are specific or generic, whether they sound current, and above all whether the tool acknowledges that its own knowledge may be out of date. Do the same for your industry question: cross-check one specific claim against a source you trust, such as a regulator's website, an industry publication, or your own experience, and then ask the tool directly, "How confident are you in this, and what should I verify before relying on it?"
The verification step is non-negotiable for factual claims. These tools are trained to produce fluent, confident text; they are not trained to know when they are wrong. The error rate on specific factual questions varies significantly by domain and by recency. For well-documented, stable areas, reliability is high. For recent developments, niche technical topics, or jurisdiction-specific regulation, errors are more common and are stated with exactly the same assurance as the parts that are correct.
Building Your Personal Observation Log
Keep the notes in a simple, repeatable shape so that comparisons across exercises and across tools actually mean something. After each exercise, capture the tool you used, the task, what worked well, what disappointed you, what surprised you, and whether you would use this tool for this task again. Four sentences per exercise is enough. A short template keeps you honest and takes almost no time to fill in:
- Tool used: which assistant, and which version or tier if you know it.
- Task: a one-line description of what you asked for.
- What worked well: what did the tool do particularly well?
- What was disappointing: what did not work as well as you expected?
- Surprises: anything unexpected, in either direction.
- Would I use this tool for this task again? Yes, no, or maybe, with one clause of reasoning.
- Tool I would recommend for this task: name it, or record that several were equally good.
Over time this log becomes a practical reference. When a colleague asks whether AI can help with something, you will have evidence rather than impressions. When a new tool launches and someone asks how it compares, you will have a baseline drawn from direct experience rather than from vendor claims. Come back to these exercises in two to three months. The tools will have changed, and more importantly so will you: experience with AI makes you a more precise observer of AI, so what surprised you the first time will be routine, and what surprises you the second time will be subtler and more informative.
What Your Observations Should Show You
After the five exercises, your notes should contain evidence for conclusions along these lines. Do not take them on trust; check each one against what you actually wrote down, and if your experience contradicts a point here, your experience is the more useful record.
- Context specificity is the primary driver of quality. Generic prompts produce generic outputs, while role, audience, constraints and format transform the same request into something usable.
- Follow-up questions are more powerful than initial prompts. The first response is rarely the best one, and a well-targeted second pass repeatedly produced the insight the first prompt missed.
- Everything needs review before use. The output was mostly right and still needed refinement, which is the normal case rather than a bad run.
- Factual claims require independent verification. Confidence of tone carries no information about accuracy, because the AI does not know what it does not know.
- AI excels at volume generation, not selection. Brainstorming and drafting are strong; deciding which of ten ideas fits your context is your job, not the tool's.
- Capability is uneven across tasks. Email and summarisation tend to feel polished, brainstorming is useful but sometimes generic, and data analysis is helpful but wants checking.
- Different tools have genuine strengths. If you ran the exercises on more than one, you probably found one better at drafting and another better at analysis. Document that and build a personal map of which you reach for when.
- It is a thinking partner, not a replacement. The best results came where you used the output as a starting point and then applied your own judgement.
Limitations These Exercises Should Have Exposed
Some limits show up consistently across tools, and running into them yourself is far more persuasive than reading a warning. There is no access to real-time information, so any question that depended on current events or recent data may have been answered from a stale picture. There is no verification of factual claims, because the model does not check its own facts before stating them. There is no understanding of your specific context, so budget constraints, organisational culture and history are invisible unless you type them in.
Three more are worth naming because they are easy to miss while the output is impressing you. There is no accountability: if the advice is wrong and you act on it, the responsibility is yours. The tools tend to be agreeable, offering diplomacy where you may have wanted the hard truth, which makes them poor at telling you your idea is bad. And they are sometimes confidently wrong, presenting incorrect information in precisely the tone they use for correct information. That last one is the reason tone can never be your quality signal, and the reason the verification habit from Exercise 5 belongs in every workflow you build from here.
Anti-Patterns
The exercises are simple enough that the ways they fail are all about how they are run rather than about what they contain.
- Reading the exercises instead of running them. The chapter is written to be executed. Reading a prompt and imagining the response produces the illusion of having learned something, and the imagined output is always tidier than the real one.
- Skipping the deliberately weak prompt. It is tempting to go straight to the detailed version because you already know it will do better. The contrast between the vague and the specific run is the actual finding; without the baseline you have an output, not an observation.
- Grading on fluency. Every response arrives polished, so it is easy to mark work as good because it reads well. The checks that matter compare the output against something outside it: the source document, the numbers in the table, a regulator's page, your own knowledge of the situation.
- Testing where you cannot tell right from wrong. Choosing a research question in an area you know nothing about feels like a fair test and teaches you nothing, because you have no way to grade the answer. Pick at least one question where you already know roughly what the response should say.
- Relying on memory instead of notes. The details fade within an hour, and what remains is a general impression that is usually more positive than the session deserved. Impressions are exactly what the log exists to replace.
- Generalising from one tool. Conclusions drawn from a single assistant describe that assistant. If you say "AI is bad at data analysis" on the strength of one run, you have made a claim your evidence does not support.
- Treating the follow-up as optional. The second prompt is where most of the learning sits, particularly the question about what the tool cannot confirm. Stopping at the first response measures the tool at its least useful.
Practice Prompts
These extend the exercise set into your own work, once you have run the five and have a log to compare against.
- Do it live. Take a real email you owe someone and run Exercise 1 on it, with your actual context. Note how much you edited before sending, and whether the edits were things you could have specified in the prompt.
- Run the same prompt twice, on two tools. Use one exercise, unchanged, on two different assistants. Record where they diverged. That comparison is the beginning of a personal map of which to reach for when.
- Audit a summary against its source. Take any summary you have generated and mark every claim that does not appear in the original document. Anything you mark is an error rather than a stylistic choice.
- Ask for the verification list. On your most recent AI-assisted factual claim, ask the tool how confident it is and what you should verify before relying on it, then actually verify one item.
- Separate observation from inference. Take a piece of analysis the tool produced and mark each sentence as either something in the data or an explanation of it. Weak outputs blend the two at the same level of confidence.
- Write the colleague version. Turn your log into one page telling a colleague which tool to use for which task and what to check afterwards. Anything you cannot support from your notes is an impression, not a finding.
Reflection
Look back over your notes and ask what surprised you, and then ask why it surprised you, because the second question is where your working model of these tools was wrong. Consider whether your expectations were too high or too low going in, and which exercise moved them furthest. Then think about the check you skipped: almost everyone skips one, usually the independent verification in Exercise 5, because the answer sounded right. What would it have cost you to be wrong about that claim in a piece of real work? Finally, ask which of the tasks you just tested you would now hand to an AI tool without a second thought, which you would hand over with review, and which you would keep entirely for yourself. Those three lists, drawn from your own evidence, are more useful than any general claim about what AI can do.
Glossary
- Baseline prompt: A deliberately underspecified first version, run so that the improvement from a detailed version can be seen rather than assumed.
- Framing: The role, purpose or perspective supplied in a prompt, which determines what the tool treats as important and therefore what a summary surfaces or omits.
- Follow-up prompt: A second instruction sent after reading the first response, used to refine, redirect, or probe what the tool cannot confirm.
- Observation log: A short written record per exercise covering the tool, the task, what worked, what disappointed, what surprised you, and whether you would use it again.
- Verification: Checking a specific claim against a trusted source outside the tool, which is non-negotiable for factual material regardless of how confident the output sounds.
- Confidently wrong: Incorrect information presented in exactly the tone used for correct information, which is why tone can never serve as a quality signal.
- Knowledge cutoff: The reason a tool has no access to real-time information, so answers about recent developments may be drawn from a stale picture.
- Extrapolation: Projecting beyond the data you supplied, as in forecasting the months after a table ends, which is speculation presented in the same register as the analysis.
Related Lessons
This chapter assumes you have a tool open and know roughly what it is. AI Tool Orientation and Understanding How AI Tools Work cover that ground, and are worth revisiting if any exercise produced a result you could not explain. Once you have the observations, two lessons turn them into practice: Critical Evaluation Framework systematises the judging you did informally here, and Quality Decision Framework covers what to do with an output once you have judged it, including when refining beats starting over.
Closing
Preethi's summary of her first attempt holds up as a description of the whole set: fast, fluent, and wrong in ways she could only catch because she knew the domain. Nothing in this chapter argues that the tools are weak. It argues that their strengths and their failures are both specific, and that specific knowledge only comes from watching them work on tasks where you can grade the result. You now have a record of that. Keep it, add to it whenever something surprises you, and treat it as the thing you reason from when someone asks whether AI can help with a task. The alternative is arguing from impressions, and impressions are exactly what fluent output is best at producing.
Key Takeaways
- Practical experience beats theoretical knowledge for building AI intuition. These exercises create the direct observation record that reading about AI cannot provide.
- Context specificity is the single biggest driver of output quality. Role, audience, format, and constraints in your prompt turn generic outputs into usable ones.
- Follow-up prompts are often more valuable than initial prompts. Iterating from a first response consistently outperforms refining the initial prompt in isolation.
- Factual claims require verification, regardless of how confident the output sounds. Cross-check anything you intend to rely on against trusted sources, and ask the tool what it thinks you should verify.
- Invented detail is an error, not a style problem. Anything in a summary that was not in the source document is a failure of the task, and it is the first thing to look for.
- AI generates volume well; you select and judge. The value of brainstorming is as a trigger for your thinking, not as a ready-made answer.
- Build a personal observation log. Written notes from direct experience are more accurate and more useful than impressions formed after the fact.
- Revisit these exercises periodically. Both the tools and your ability to observe them change over time, so periodic re-testing is more informative than a single evaluation.
Frequently Asked Questions
Which tool should I use for these exercises? Any assistant you can type into will work, whether that is a general-purpose chat tool, one built into your productivity suite, or whatever your organisation has approved. The exercises are written to be tool-neutral. If you have access to more than one, running an exercise on two of them is worth the extra few minutes, because the differences are part of what you are trying to learn.
Do I have to do them in order? The order moves from the most forgiving task to the least, ending with research, where confident wrong answers do the most damage. If you only have time for some of them, keep Exercise 5, because the verification habit it builds is the one that protects real work.
How long should this take? Plan for about forty to fifty minutes across the set. Rushing defeats the purpose, since the value is in noticing details rather than in completing the list, and most of the noticing happens while you compare a follow-up against the response that preceded it.
Can I use real work material? Only material you are permitted to share with the tool you are using. Several exercises invite you to paste a document or a table, and the safe choice is something you would be comfortable sending outside your organisation, or a sample like the one provided here. Where you are unsure, use the sample and run the exercise on your own material later, under whatever rules apply to you.
What if my results do not match what the chapter predicts? Then your notes are the more accurate record and should be trusted over the general description. Tools differ, they change, and the tasks you bring are not the tasks used here. A result that defies your expectation is the most informative outcome available, which is why the chapter asks you not to skip exercises whose outcome you think you can guess.
Is it worth repeating this later? Yes, and two to three months is a reasonable gap. The tools will have changed, and so will your ability to observe them. What surprised you the first time tends to be routine on the second pass, and the new surprises are subtler and more useful, which is exactly the progression the log is designed to capture.
Skill.re