Reasoning Models in Creative: Territories Before Executions

Reasoning models change creative development by making exploration cheap and critique rigorous. Instead of asking for finished ads, teams should ask models to map creative territories, test each against audience evidence and brand rules, and explain trade-offs. Humans then choose the territory and craft the work. The model widens the field; people decide what is worth making.
Why do most AI creative workflows disappoint?
Because they skip to the end. A team types "write five Instagram ads for our new range" and receives five competent, interchangeable ads. The problem is not the model's writing ability; it is that nobody asked it to think about the idea. Creative quality comes from the choice of territory, the insight behind it and the specific human truth it dramatises. Execution is the last ten per cent.
Reasoning-era models such as GPT-5, which OpenAI built with a deeper thinking mode for complex problems, are well suited to the earlier, harder ninety per cent: exploring territories, articulating the insight behind each, and critiquing them against criteria. At GPT5 Marketing, an independent frontier AI strategy studio, we redesign creative workflows to use them there.
What does a reasoning-led creative workflow look like?
- Ground: load the brand memory pack and current audience signal, including the exact language people use about the tension.
- Map: ask the model for six to ten distinct creative territories, each with the insight it rests on, the evidence behind that insight and the emotional register.
- Critique: ask the model to score each territory against written criteria: distinctiveness in category, fit with brand, strength of evidence, production feasibility.
- Choose: creative leads pick one or two territories, often combining or rewriting.
- Craft: writers and designers develop the work, using fast-mode models for variants and adaptations.
- Pre-test: check clickability and resonance before spend.
For the pre-test step, SOMIN, an AI audience-research platform, offers pre-spend creative prediction in its paid-media product, which pairs well with the territory work upstream.
How do you stop AI creative from sounding the same as everyone else?
Sameness comes from shared defaults. If every brand asks a model for "warm, authentic, relatable" copy with no specific evidence, the outputs converge. Three habits break the pattern:
- Feed real audience language. Specific phrases from real posts pull the model away from category clichés.
- Demand distinctiveness explicitly. Include competitor creative in context and ask the model to identify what territory competitors already own and avoid it.
- Ask for the uncomfortable option. Request one territory that the brand would find risky. It is often the most interesting starting point.
The Pepsi India case study is worth reading for how audience-led thinking shapes creative choices in a crowded, fast-moving category.
What is the role of multimodal input?
OpenAI reported strong GPT-5 results on multimodal benchmarks covering visual and video-based reasoning. For creative teams, this means a model can look at a mood board, a competitor's ad or a storyboard and reason about it alongside text. Useful applications include critiquing a key visual against brand rules, comparing your shelf presence with competitors' or describing what a storyboard communicates to someone who has not read the script. Treat these as structured second opinions, not verdicts.
Definitions
- Creative territory: a distinct strategic and emotional direction for creative work, defined before execution.
- Critique pass: a reasoning step in which a model evaluates options against written criteria and explains its scoring.
- Pre-spend testing: predicting creative performance before media money is committed.
Where do creative people fit?
Right at the centre, with a different job description. Creative leads spend less time generating first ideas from a blank page and more time choosing, combining and sharpening. Writers spend less time on the twentieth variant and more on the line that carries the campaign. This is the "enhance, not replace" pattern in practice, and it reflects a tension we hear constantly in listening data from marketers worried about their roles. We explore that tension in a dedicated article.
Agencies building fully agentic pipelines, such as our sister business AgentC, make the same point from the other direction: agents can run research to creative to optimisation, but human strategists own judgment.
A worked example
A home-insurance brand wants to reach renters. The grounded signal shows a recurring tension: renters feel insurance is "for homeowners" and assume their belongings are not worth protecting, until something happens to a laptop. The model maps eight territories, from "your stuff is worth more than you think" to a deliberately risky "the landlord's insurance does not cover you". The critique pass flags the second as most distinctive and best evidenced, but legally sensitive. The creative lead chooses it, works with legal on precise claims, and the team crafts executions around specific moments: moving day, the first flood warning, a stolen bike.
None of that required the model to write the final ad. It required the model to think well about which ad was worth writing.
One habit worth adopting immediately: keep a territory log. Every territory the model proposes, and every one the team rejects, goes into a simple record with the reason. Over a year, that log becomes a map of what your brand has considered and why, which is invaluable when a new agency, a new chief marketing officer or a new model joins the process. It also stops teams rediscovering the same rejected idea every second quarter.
Checklist for your next brief
- Did you ask for territories before executions?
- Is each territory linked to real audience evidence?
- Did the model see competitor creative?
- Are the critique criteria written down?
- Is there a pre-spend test before media money moves?
Frequently asked questions
How should creative teams use reasoning models?
Ask for distinct creative territories with evidence and trade-offs, critique them against written criteria, then let people choose and craft the work.
How do you avoid AI creative that sounds the same?
Feed real audience language, include competitor creative in context, demand distinctiveness explicitly and ask for at least one uncomfortable territory.
Can reasoning models judge visual creative?
GPT-5 performs strongly on multimodal benchmarks, so it can critique key visuals or storyboards against brand rules. Treat this as a structured second opinion, not a verdict.
Start a frontier sprint
An independent frontier AI strategy studio for CMOs. We rebuild strategy, creative, research and measurement workflows for the reasoning era, with SOMIN as the evidence layer underneath.
Email ask@gpt5.marketing →

