|

AI Marketing Research: Creativity, Personas & Agents

If your team is picking AI tools on vibes — one model for copy, another for brainstorming, whichever chatbot persona sounded friendliest in the demo — you're not alone. Until recently, there wasn't much rigorous AI marketing research to guide those calls.

Three new preprints released together start to change that. One introduces the largest systematic benchmark of AI creativity to date. Another proposes a framework for chatbots that flex their role and personality intensity by context. The third simulates what happens when AI agents compete for work in an automated marketplace.

All three are preprints, so treat the findings as strong hypotheses rather than settled science. But the throughline is clear: default settings — one model for every creative job, one cheerful persona for every conversation, one generic prompt for every agent task — are leaving performance on the table.

This piece is written for marketing leaders, agency operators, and founders deciding how to deploy AI across creative work, customer conversations, and automated workflows. Here's what each paper suggests, what it doesn't prove, and what to test next.

Quick Takeaway

  • Adding 'be creative' to prompts boosted AI creative output more than reasoning modes in one benchmark study.
  • AI models share a general creativity factor, but domain leaders differ — pick the model per task type.
  • Medium-intensity chatbot personas tend to outperform flat or overly enthusiastic ones on trust and likeability.
  • AI agents that self-assess, track competitors, and plan ahead captured 1.5x more market share in simulation.
  • All three papers are unreviewed preprints — directional guidance, not confirmed laws.

What This Research Means for Marketers

Each paper points at the same underlying issue: AI performance is highly sensitive to how you set it up, and most teams are using defaults. Small changes — a two-word prompt addition, a mid-intensity persona instead of maximum-friendly, a scaffolding layer that tells an agent to plan ahead — meaningfully change outputs in these studies.

The practical implication isn't 'buy a new tool.' It's 'audit your prompts, personas, and agent instructions.' The teams getting real leverage from AI aren't the ones with the most subscriptions — they're the ones treating configuration as a discipline.

Papers Covered

Paper 1: AGC-Bench: Measuring Artificial General Creativity

  • Source / venue: arXiv (preprint)
  • Link: unknown
  • Source type: Systematic review and benchmarking study
  • Method: PRISMA-compliant review of 3,101 papers identified 497 AI creativity benchmarks; 78 were onboarded into a HELM-style evaluation harness across six domains. 83 LLMs were evaluated. LLM-as-judge bias was corrected using Judge Response Theory, and a fine-tuned scoring model (AGC-Judge) was created. Human comparison was conducted on a five-task subset.
  • Sample: 83 frontier and open-weight LLMs across 78 benchmarks (67 text-only, 11 multimodal); 48,299 calibrated ratings; human comparison on 5 tasks with an unspecified number of participants.
  • Main finding: A single 'creativity factor' explained 81.5% of variance in creative performance across 83 models. Prompting a model to 'be creative' improved output more than enabling step-by-step reasoning. Domain leaders differed — Claude models led on narrative and figurative language, GPT-5.4 led on brainstorming.
  • Evidence strength: Preprint, not peer reviewed; large-scale benchmarking but text-heavy and human comparison limited to 5 tasks.
  • Limitation: Only 78 of 497 identified benchmarks were run; multimodal creativity is underrepresented; human comparison sample is thin; AGC-Judge trained on ratings from only three frontier LLM judges.
  • Practical implication: Add 'be creative' to prompts for copy and concept work as a low-cost quality lift, and select models by task type rather than defaulting to one for all creative jobs.

Paper 2: Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework

  • Source / venue: arXiv (presented at AAAI-2026 Bridge Program: Bridging AI and Behavior Change)
  • Link: unknown
  • Source type: Conceptual framework / position paper
  • Method: Synthesis of two existing empirical research streams — metaphorical persona design in voice interfaces and personality expression intensity in LLM-based agents — into a proposed unified design framework. No new empirical data collected.
  • Sample: unknown
  • Main finding: Chatbots that shift role by context (coach, tutor, tool) tend to be received better than static bots. Medium personality intensity generally outperforms flat or highly enthusiastic settings on trust and likeability. Human-like personas can backfire when they raise unrealistic expectations. The paper proposes tuning both role and intensity together.
  • Evidence strength: Preprint conceptual framework; no direct empirical test in this paper. Inferences drawn from prior studies whose findings are context-dependent.
  • Limitation: Framework is not empirically validated here; the 'medium is best' pattern did not hold in every prior study cited; implementation details for context detection are sketched but not fully specified.
  • Practical implication: Write context-specific persona prompts (support, sales, celebration) rather than one generic assistant voice, and test medium-intensity personalities against maximally friendly defaults.

Paper 3: When AI Agents Compete for Jobs: Strategic Capabilities and Economic Dynamics of AI Labour Markets

  • Source / venue: arXiv (preprint)
  • Link: unknown
  • Source type: Simulation / computational experiment
  • Method: Authors built a gig-economy testbed called AI-Work in which LLM agents bid on jobs, invest in training, manage reputation, and adapt strategies under partial observability across discrete rounds. Experiments varied platform rules, prompting strategies, and market structures.
  • Sample: Simulated GPT-class LLM agents across multiple experimental conditions. No human participants. Exact agent counts not specified in available text.
  • Main finding: Agents with self-awareness, competitor awareness, and multi-step planning outperformed baseline agents. A prompting scaffold adding these three elements captured 1.5x more market share using the same base model. Public pricing triggered race-to-the-bottom dynamics; hidden pricing shifted competition toward skill investment. Broader task variety reduced winner-takes-all concentration.
  • Evidence strength: Preprint simulation only; stylized market, not empirical labour data; directional not predictive.
  • Limitation: Simulated environment inspired by gig platforms may not generalize to enterprise AI procurement or ad tech marketplaces; full experimental parameters not disclosed in available text.
  • Practical implication: When deploying AI agents for marketing tasks, add scaffolding prompts covering self-assessment, competitive context, and planning — and diversify the tasks assigned across agents rather than concentrating on a single 'best' agent.

Plain-English Payoff

Small configuration changes — a two-word prompt addition, a calibrated chatbot persona, a scaffolding prompt for agents — moved outcomes meaningfully in each of these studies. The teams getting the most from AI aren't buying more tools; they're treating prompts, personas, and agent instructions as design surfaces worth testing.

Money Move

The clearest opportunity sits at the intersection of all three papers: a prompt optimization and persona management layer for marketing teams deploying AI at scale. A lightweight system that routes each task to the right model, applies the right personality mode for the context, and injects strategic scaffolding for agent workflows — rather than running everything through one generic assistant. Agencies managing high volumes of AI-generated content and client-facing chatbots are the natural first customer.

Evidence Check

  • All three papers are preprints and have not completed peer review; findings may change.
  • Paper 1 (AGC-Bench) is a large systematic benchmark but text-heavy; multimodal creativity is underrepresented and human comparison is thin.
  • Paper 2 (Fluid Personality Framework) is a conceptual synthesis with no new empirical test — the design pattern is inferred from prior work.
  • Paper 3 (AI-Work) is a simulation, not a real labour market; results are directional and illustrative, not predictive.
  • Do not overclaim causal effects on real customer behaviour from a simulation or a benchmark of model outputs alone.
  • Full parameter details (sample sizes, exact model lists) are not disclosed in every case — treat specific numbers as provisional.

What to Test Next

  • Action step. Run the same creative brief through your AI tool twice — once with your standard prompt, once with 'be creative' added — and have a colleague blind-rate the outputs to see whether the lift replicates for your use case.
  • Action step. Audit your customer-facing chatbot for persona-context mismatches. Identify at least one high-stakes flow (like complaint handling) that currently uses a uniformly cheerful voice, and draft a calmer, medium-intensity variant to A/B test.
  • Action step. If you're deploying AI agents for any autonomous task, add a scaffolding prompt covering self-assessment, competitive awareness, and multi-step planning, and measure completion quality against your baseline prompt.
  • Action step. Build a simple internal reference of which AI models perform best for which creative task types in your workflows, and stop defaulting to one model for all creative work.

How This Connects to AI Marketing Strategy

Across the Radar's coverage, the same pattern keeps surfacing: AI's business value is less about which model you pick and more about how you configure the space around it — prompts, personas, task routing, evaluation. These three papers formalize that intuition across three surfaces at once: creative generation, conversational interfaces, and autonomous agents.

For marketing teams, that reframes the strategy conversation. The competitive edge isn't in access to frontier models — everyone has that. It's in the operational discipline of testing prompts, calibrating personas, and scaffolding agents against real business outcomes. That's a durable capability, and it's what separates AI-native marketing operations from teams that are still running everything through one default assistant.

FAQ

What is AGC-Bench and why does it matter for marketers?

AGC-Bench is a systematic benchmark that evaluated 83 AI models across 78 creativity tasks in six domains. It matters because it gives marketers the first standardized way to compare models on creative work like copywriting, brainstorming, and figurative language — replacing anecdotal 'which model is best' debates with domain-specific evidence. It's a preprint, so treat findings as strong hypotheses.

Does adding 'be creative' to an AI prompt actually work?

In the AGC-Bench study, adding a simple 'be creative' instruction improved creative output more than turning on step-by-step reasoning modes. It's a cheap, testable change, but the finding comes from a preprint benchmark of model outputs — not from a controlled customer experiment. Run your own blind comparison before treating it as a rule.

What personality should my AI chatbot use?

The Fluid Personality Framework paper suggests a medium-intensity personality tends to outperform both flat and overly enthusiastic settings on trust and likeability, and that the bot's role should shift with context (e.g., calmer for complaints, warmer for onboarding). It's a conceptual framework rather than a validated test, so A/B test it against your current defaults.

Are AI agents ready to run marketing campaigns autonomously?

Not without careful design. The AI-Work simulation showed that agents with basic scaffolding — self-awareness, competitor awareness, planning — dramatically outperformed baseline agents using the same underlying model. Marketers deploying agent workflows should invest in prompt scaffolding and diversify tasks across agents to avoid fragility.

How can small businesses use these findings?

The most accessible move is prompt-level: add 'be creative' to creative tasks, write context-specific chatbot personas rather than one generic voice, and give any AI agent explicit instructions to plan ahead. None of these require new tools or budget — just a more disciplined approach to configuration.

Are these findings peer reviewed?

No. All three papers are preprints on arXiv and have not completed peer review. They're worth acting on as testable hypotheses, but avoid citing them as settled science in client work or major strategic decisions until validated further.

What's the biggest risk in acting on this research?

Overclaiming. A benchmark of model outputs isn't the same as a study of customer behaviour, and a simulated agent marketplace isn't a real one. Use these papers to shape internal tests, not to justify sweeping claims about ROI, creativity, or agent performance in production.

Listen to the Episode

Listen on Buzzsprout

Sources and Further Reading

  1. AGC-Bench: Measuring Artificial General Creativity
  2. Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework
  3. When AI Agents Compete for Jobs: Strategic Capabilities and Economic Dynamics of AI Labour Markets

Related Big Plans Media:

About Big Plans Media

Big Plans Media helps marketers, educators, entrepreneurs, consultants, and business leaders translate AI marketing research into practical strategy. AI & Marketing Research Radar is produced by Big Plans Media and hosted by Evita, an AI-generated research briefing avatar trained on Dr. Eva Wolf's research framework.



Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *