AI Ads That Don't Look AI-Generated: The Evidence and the Recipe

Updated September 11, 2026·35 min read·Business
TL;DR

AI-generated ads match human-made ads on click-through once campaign controls go in, and the thing that actually costs you clicks is looking AI-generated, whichever way the ad was made.

A working paper covering 369 million matched ad impressions finds AI-generated and human-made ads at parity on click-through once campaign controls go in, and underneath that parity sits a pattern the authors flag but never model causally: ads that read as AI underperform whichever way they were made, so the only combination beating the human baseline is AI creative nobody clocks as AI, and human ads that happen to look synthetic come last of the four.

The cue list, the vendor teardown, the cost tables and the legal section are linked; the de-slop pass, the prompt grammar and the QA checklist come out of pipelines I have shipped and never A/B tested against click-through.

Do AI-generated ads work?

In matched field data, yes, at parity. Exner, Hartmann, Netzer and Zhang ran a quasi-experiment covering more than 16.4 billion impressions and 116 million clicks across 305,121 display and native ads, then narrowed to 4,633 sibling ads matched on advertiser, objective and landing page and differing only in whether the image was AI-generated. There, AI images returned a raw click-through of 0.76% against 0.65% for human images across 369,533,326 impressions, and once campaign fixed effects go in the gap stops being significant, which the authors summarise as AI ads performing comparably at far lower cost.

These are static images on display and native inventory, so the cues should transfer to vertical social video while the magnitudes should not be assumed to. Native inventory also runs a weak creative baseline, and the largest field RCTs on generative AI in retail workflows found sales treatment effects from 0% to 16.3% concentrated where the baseline was poor, so if your creative is already good, plan for the bottom of that range.

The four cells that matter

Human raters scored 1,751 of the images blind for perceived artificiality, and people are bad at it: 24.87% of genuinely human images landed in the likely-or-definitely-AI bucket, and 58.92% of genuinely AI images in the not-sure-or-likely-or-definitely-human bucket. Crossing origin against perception gives four cells.

The image wasIt read asMean CTR
AI-generatedHuman-made, or unclear0.79%
Human-madeHuman-made, or unclear0.67%
AI-generatedAI-generated0.62%
Human-madeAI-generated0.55%

Unadjusted cell means, no cell sizes or confidence intervals published, perceived artificiality never randomly assigned, and the paper is not peer reviewed as of September 2026. The paper's own term for the top row is images that "disguise their origin".

Treat that ordering as a hypothesis about a production cue, since you cannot choose your cell: at the study's detection rates a batch of AI creative blends to roughly 0.72% against roughly 0.64% for human creative, the same raw gap the controls cannot distinguish from zero. The bottom row should bother you, because nothing generative touched it and the ad still comes last.

Why HBR and Ipsos found the opposite

In May 2026, Ipsos and two Syracuse Newhouse researchers, Adam Peruta and Carrie Riby, tested 20 ads across 10 brands with 3,000 US consumers and found human-made ads over-indexing the benchmark by 11 points while AI-made ads under-indexed by five, which Harvard Business Review wrote up in September 2026 under a headline saying AI ads perform worse even when customers cannot tell them apart. Almost every page I can find cites one study and pretends the other does not exist.

They measure different things. Peruta describes ads that "started from the same brief. The only thing that changed was who made them," setting produced campaigns from ten large brands against reconstructions built for the study, so the variable under test is craft tier wearing an AI label, scored as predicted effect on a survey panel instead of delivered clicks in a live auction. They agree where it counts: only about 13% of Ipsos viewers were confident an AI ad was AI, and the same proportion suspected the human ads. Kapoor and Kumar in Marketing Science sit between the two, finding generative personalised video lifting engagement 6 to 9 percentage points over a personalised image and a generic video, confounded with personalisation, as does Harvard Business School's summary of the Exner work. Match a competitor's in-feed creative and expect parity; rebuild a finished national campaign from a brief and expect to lose, because there the variable is your ability to direct.

The numbers everyone repeats, and where they came from

If you have searched this topic you have seen the cluster: 12% to 23% higher CTR on Meta, 21% lower CPA, 3.4x ROAS against 3.1x, 89% cheaper per asset, 14.3 variants against 3.7, on dozens of pages with no attribution.

One vendor blog, five sources, four of them vendors

Every one of those figures I have traced runs back to a single content-marketing page by Lapis, an AI ad generation company, whose five cited sources are Taboola's generative AI study at roughly 500 million impressions, Soku.ai at 2,500+ campaigns, Social Operator at 1,200+, Genesis/Synthesia at 800+ and DigitalApplied at 1,500+: four vendors and an SEO publisher.

Taboola is the one academically grounded name there, and the paper written on Taboola's data is the Exner work above, while what Lapis reports as the Taboola study is a 500-million-impression headline test, a thirtieth of the size and about text instead of imagery. Lapis carries a caveat nobody repeats downstream, that its figures come from "real-world campaigns with real budgets" rather than a controlled A/B test, so "the data is messy, as real data always is."

Meta's own lifts are single digits

The platform with the most to gain from AI creative publishes three figures on its own Advantage+ creative page, as of 10 September 2026, with no methodology, sample or period attached, on a page it edits without version history:

  • 2% to 3% lift in conversions from background generation on catalog ads
  • 2% lift in conversions on Facebook Reels from video expansion
  • 13% more conversions using related media

The often-quoted 22% ROAS number comes from Meta's engineering blog in December 2024, where the full sentence changes it: "when advertisers who did not previously use Advantage+ creative turned on its AI-driven targeting features, they experienced a 22% increase in ROAS from our ads." That is a new-adopter figure about targeting, and the same post says Meta estimates a 7% conversion increase for image generation, where estimate is doing real work. I would believe the single digits, because when the first-party source with every incentive to inflate publishes 2% to 3% for background generation and 13% at the top end, a third-party page promising a fifth or a half more revenue from AI creative is measuring something else and has not told you what.

A ten-second test you can run on any AI-ad statistic

Ask who published it and whether they sell the thing the number flatters, ask for an n, a window and a method, and follow the citation chain to its end, since a vendor at the end is a rumour with footnotes.

Run the same test on my own tables and it draws blood. Motion sells creative analytics whose value proposition is that testing volume matters, and the volume argument below is sourced entirely to Motion, whose data is not independently auditable. Amplified Intelligence, VCCP and Dentsu sell attention measurement or media planning, and none publishes an auditable method, including the r = .82 correlation I quote without an n. invideo sells generation credits and supplies the yield figures and per-ad cost, Icon sells human UGC and supplies the comparison price, and every one of them has a directional interest in the number I am borrowing.

The study most readers arrive already believing needs its date attached. NIQ's EEG and eye-tracking research, 2,000+ participants with about 150 on EEG, found AI ads rated more annoying, boring and confusing with weaker memory activation. It is a press release from 12 December 2024 with no per-format numbers and no stimuli, and every current video model postdates it, so what it measures is reaction to 2024-vintage AI creative.

How to make AI ads that don't look AI-generated: the de-slop pass

AI slop is creative carrying the cues consumers read as machine-made, worth 0.12 to 0.17 percentage points of click-through in the matched data, unadjusted, whichever way the image was made. Every "make your AI ads look real" guide I have read recommends shallow depth of field, an 85mm look, soft window light and a cinematic grade, which needs a second look against the three features that moved perceived artificiality:

  • Intense colour saturation signals AI generation to consumers, one cue out of several and the one with a clean production fix.
  • Images displaying text are perceived as more likely to be AI-generated, which the authors attribute to diffusion models rendering text badly. Barely discussed anywhere, and the consequence is direct: do not bake copy into the frame at generation time.
  • High aesthetic scores and medium-to-large faces are associated with lower perceived artificiality, and the AI images in the dataset had larger faces and better aesthetics, which consumers read as evidence of a human.

What the de-slop pass cannot do

The pass targets human perception and does nothing about machine detection, which is where platform labels come from, so planning around beating raters and getting labelled anyway buys the worst of both.

Meta labels ad images itself when created or significantly edited with its own or third-party generative AI tools, TikTok reads C2PA Content Credentials on uploads and says it can instantly recognise and label AI content from that metadata whatever the creator declared, and YouTube auto-applies its label to content carrying C2PA metadata. Google's SynthID watermark is designed to survive cropping, added filters, frame-rate changes and lossy compression, a fair description of what a realism pass does to a frame, and ByteDance's Dreamina and Seedance surfaces ship output with invisible watermarks, C2PA Content Credentials and visible AI labels, so check what your account attaches before committing creative.

If you are planning around a label, budget the 31.5% click-through cost the NYU Stern work below puts on one. The good news is buried in the policies, since YouTube states that disclosing AI content "won't limit a video's audience or impact its eligibility to earn money," and TikTok says the same about distribution, so a label costs you reader response while delivery carries on.

The recipe, as preserve, add, avoid

I run this image-to-image pass on every storyboard before it becomes video, changing micro-realism only.

Preserve, explicitly:

  • each face's shape, width and proportions, one to one
  • framing, composition and camera distance
  • pose, gesture, product identity and label placement

Image models love to slim faces on a second pass, and a re-sculpted face is its own tell.

Add:

  • pore-level skin with fine vellus hair
  • real material detail in fabric, plastic, metal and paper
  • even daytime light with gentle highlight roll-off
  • faint true sensor noise in shadow
  • deep focus, with the background sharp
  • the look of a flat, unedited phone photo

Avoid:

  • waxy or poreless skin and beauty-filter smoothing
  • over-saturation, HDR glow, bloom and halos
  • oversharpening and the teal-orange grade
  • shallow depth of field, bokeh, the cinematic DSLR look generally
  • any text or watermark in the frame

What the evidence does not say, and what I do anyway

Only desaturating and keeping text out of frame come from the measurement, and part of the rest contradicts the paper, since blur in its feature table is small and not significant while higher aesthetic quality predicts lower perceived artificiality. I strip the grade anyway because a shallow-focus 85mm look pushes the face smaller in frame, face size is the strongest human-reading cue the study does measure, and the register consumers read as human in UGC is amateur. Untested hypothesis, and I will call it one: nothing links this pass to click-through in any A/B test I can point at. The rest of the phone-camera brief has the same status, a 23mm-wide look with mild edge distortion, one autofocus adjustment mid-clip, real contact shadows, and hair and fabric that move. Text belongs in post, so burn captions from a word-level transcript after the render, which fixes the text cue and the gibberish typography at once, and the prompt-side version is in the Seedance 2.5 production guide.

How long does someone have to watch before an ad works?

Roughly three seconds, unless you already own distinctive brand assets, in which case the published threshold drops to about 1.5. Karen Nelson-Field's work with Amplified Intelligence, published through WARC, reports r = .82 between active attention seconds and short-term advertising strength, no sample size published, with memory retention beginning at roughly three seconds and each further second buying about three days in memory.

VCCP Media and Nelson-Field then ran 20,000+ views of 72 digital video ads in a good-twin/bad-twin design and found 1.5 seconds of genuine attention enough to encode a memory, with well-branded ads 2.5x more effective in low-attention environments, against the industry assumption that 85% of digital placements get under 2.5 seconds. She states the caveat plainly in her Creative Salon interview: 1.5 seconds holds only for a brand that already owns distinctive assets, so a three-month-old DTC brand is on the three-second clock.

Dentsu Japan's eye-tracking study of 8,000+ mobile users across 45+ creative variations gives a recall curve that is not monotonic:

Viewing durationBrand recall
Under 1 second21%
1 to 3 seconds35%
3 to 6 seconds50%
Over 10 seconds41%

The 3-to-6-second band beating the 10+ band is probably a format confound, since long views skew toward forced pre-roll, so treat three to six seconds of earned attention as the target.

What eye-tracking says about cuts and centred faces

Ye and Wedel's ViASNet paper built a video-ad saliency model on 151 ads, each viewed by 20 participants on a 100Hz eye tracker. Frames with low gaze entropy and high engagement tend to have a person, a face or the product in the centre, high-entropy low-engagement frames spread objects or text across the frame, and entropy jumps at every cut as people reorient. The centring half agrees with the Exner finding on medium-to-large faces, though the stimuli are Dutch television commercials, so the numbers may not travel. It also cuts against the cut-every-second orthodoxy, including my own eight-cut UGC structure: eight cuts because the story has eight beats is defensible, eight because cutting feels energetic is a tax.

The sound-off myth, traced to its source

"85% of Facebook video is watched without sound" is still in 2026 agency decks, and it comes from a 2016 Digiday article citing self-reported numbers from three publishers, one of whom gave a range of 50% to 80%. It was never Meta research, no methodology was published, and it predates Reels entirely. What holds up is weaker: Kantar's work with TikTok finds 88% of TikTok users say sound is essential to the experience, an attitudinal self-report, while Meta's guidance for Facebook feed video lists captions and sound as optional but recommended. Build for legible without sound and better with it.

What is a good hook rate, and when should you kill an ad?

Hook rate is 3-second video plays divided by impressions, thumbstop ratio is the same thing renamed, and no platform publishes an official benchmark, so every figure in circulation is a practitioner aggregate. The most transparent set, Ad Library's, puts healthy cold-traffic hook rate at 25% to 35%, best-in-class above 45% and under 15% as a kill signal in feed, with 15-second hold rate healthy at 15% to 25%, kill below 10% and best-in-class 30%+. Hold rate has two competing definitions, 15-second plays / 3-second plays and Meta's `video_p75_watched / video_3sec_watched`, so say which you mean first, and expect others to disagree: Sepia Lab's carries a widely repeated TikTok average of 30.7% traced back to eleven accounts, a sample you would send an intern back to redo. Those ranges also assume a denominator most accounts do not have, since Advantage+ and Audience Network placement mixes behave nothing like in-feed Reels, so break out by placement, state the audience, and set an impression floor, because a few hundred impressions of noise will happily produce a 9% hook rate.

The correction that matters more than the ranges is that Meta and TikTok hook rates measure different things. Meta's denominator is 3-second video plays, while TikTok publishes 2-second and 6-second video views, where a 6-second view also counts a full play under six seconds or any engagement in the first six, so a deck comparing them is comparing nothing. TikTok's guidance says to prioritise the hook in the first six seconds and the proposition in the first three, run 5 to 10 words per second of on-screen text, and test 3 to 5 creatives across 3 to 5 ad groups, putting its hook window at twice the three seconds Meta buyers plan around.

The diagnostic grid

Hook and hold diagnose most of what goes wrong between them, on my own unmeasured operating logic:

HookHoldWhat it meansWhat to change
LowLowWrong angle, format or audienceReplace the ad, do not iterate it
HighLowThe opener over-promises and the body does not pay it offRewrite seconds 5 to 15, keep the hook
LowHighPeople love it once they start and will not startRebuild only the first 1.5 seconds
HighHighWorkingScale it, then test variations

Past those two, strong hook and hold with weak CTR indicts the offer or the call to action, strong CTR with weak conversion indicts the landing page or the price, and if every concept converges on a similar bad CPA, no volume of creative fixes it.

Two statistical traps that void your result

Unequal budgets make CPA comparisons invalid, because Meta allocates spend by estimated action rate, so a variant given more budget can show a worse raw CPA precisely because it reached further into a larger, less efficient audience. Use Meta's split-test tool, which randomises at the user level, compare variants at similar realised spend, and run a holdout or lift test on anything you intend to scale, because no CPA comparison across variants establishes incrementality however the budget was allocated. No platform documents that allocation effect as a named phenomenon, so treat it as reasoning about the auction.

Then Twyman's Law. In Microsoft's analysis of thousands of controlled experiments, Dmitriev and colleagues document how often a surprising result is really an instrumentation problem, and the bias to watch in your team is that people scrutinise surprising bad results and accept surprising good ones. The same paper documents novelty effects, large early gains that vanish by the second exposure, so segment by day and by visit, because a creative that wins on day two and reverts by day nine was novelty.

How many ad creatives should you test per week?

More than most accounts do. Motion's Creative Benchmarks 2026 is the largest public creative dataset, $1.29bn of realised Meta spend, 578,750 creatives, 6,015 accounts between 1 September 2025 and 1 January 2026, and across its spend tiers weekly volume rises about sevenfold while hit rate roughly doubles. A winner takes at least ten times the account's median creative spend and at least $500 absolute, and the report is clear that it measures spend concentration rather than business outcomes, so a winner by spend can be ROAS-negative.

Monthly spendCreatives per weekAverage hit rateWinners per month (arithmetic)
Micro (under $10K)2.84.0%0.5
Small ($10K to $50K)4.16.4%1.1
Medium ($50K to $200K)6.68.1%2.3
Large ($200K to $1M)11.28.6%4.1
Enterprise ($1M+)18.88.8%7.1
Note

Motion's own data, 1 September 2025 to 1 January 2026, with no refresh published since. The last column is my multiplication of Motion's first two.

That last column is where you should get suspicious of me, because Motion separately counts winners directly and gets 0.2 a month for the average Small account and 0.5 for its top quartile at 8.0 creatives a week, against my multiplication's 1.1. A five-fold disagreement inside one report means the hit rate is not the thing generating winners, and the counted number is the one to plan on. Two things also soften the tier correlation: the winner threshold is 10x the account's own median creative spend, so spreading a fixed budget across more creatives lowers the bar in absolute terms, and the top two tiers show weekly volume nearly doubling for a 0.2-point hit-rate gain, which is what concept-quality decay looks like. Winners scale with volume up to a point, and Motion's tiers put that point well before enterprise volume.

The trap the report sets up is still the most useful thing in it. Account A launches 50 creatives at a 10% hit rate and gets five winners, Account B launches five at 20% and gets one. That example assumes hit rate is invariant to volume, which is exactly what is in question, but it names the real failure: a high hit rate can mean strong briefing judgment or conservative testing, and you cannot tell which without reading it against volume.

Creative throughput matters because creative does. Westwood One, summarising NCSolutions and Nielsen, puts creative at 49% of sales lift against targeting's 11% while the marketers surveyed guessed 20% and 24%, and an earlier 2017 Nielsen study of around 500 campaigns landed close to the same split. Read 49% as directional, since buyers killing losers continuously suppress observational creative effects. This is the only argument for generative video I will defend: I am buying a seat in the volume column, because that is where winners come from.

Concepts versus variations

A concept is a different angle, protagonist, value proposition or visual language, and it answers whether the idea resonates at all; a variation iterates inside a proven one. Running five variations of an unproven concept is the commonest way to burn a test budget, because all you learn is which shade of a bad idea is least bad, so execution testing only earns budget once a concept has shown traction.

Structural defaults, practitioner consensus with no published methodology behind it, from Rocketship HQ's testing framework and campaign structure guide: keep testing campaigns separate from scaling, use ABO so variants are not starved of data, run 3 to 5 variants per ad set at 2 to 3 conversions per ad set per day, and want 30 to 50 conversions per variant before calling anything, over at least seven days because day-of-week effects swing conversion badly. Creative also burns out faster than it used to, with one vendor report relaying another vendor's analysis putting TikTok creative at roughly half its effectiveness in 72 hours, down from 120 hours in 2024, two vendors deep with no method published, though the direction matches what every buyer I know describes.

How much does an AI video ad cost?

Compute for a finished 30-second ad runs roughly $10 to $100 depending on the model and the generations you burn per keeper, while a fully loaded 60-second ad is closer to $230. Prices move monthly, so re-check this section before you quote a client.

Compute per finished 30 seconds

Routing alone can double a project. In a survey verified against provider pages on 22 August 2026, Seedance 2.5 at 720p ran $0.2312 per output second on the official ModelArk surface and on Replicate, $0.20 on WaveSpeed Turbo and $0.4730 on fal, about 2.4x from cheapest route to dearest.

Model$/secGenerations per keeperOne 30s ad
MiniMax H3 Max$0.045$10.00
Veo 3.1 Lite$0.085$20.00
Kling 3.0 Turbo 1080p$0.144$28.00
Seedance 2.0$0.1514$30.20
Wan 3.0$0.204$40.00
Seedance 2.5 720p$0.23124$46.24
Veo 3.1 Standard$0.405$100.00

Modelled from list prices as checked on 10 September 2026, assuming ten usable shots of five seconds, with the generations-per-keeper column doing all the work. Veo 3.1 Standard at $0.40 covers 720p and 1080p, with 4K priced separately.

An independent cross-check using three attempts and 70% edit retention lands on the same ordering at roughly half the absolute figures, so the ordering is stable and the levels are not. Price is the least interesting axis to pick a model on, the case I make in my head-to-head of the video models for ad work.

Reroll waste: the yields nobody puts in a pitch deck

Every list-price cost model assumes generations you keep, and yield tracks shot difficulty more than model choice:

Shot typeGenerations per keeperYield
Simple statics and product holds1 to 250% to 100%
Mixed shots with movement3 to 520% to 33%
Hands, lip-sync, walking, multi-subject6 to 1010% to 17%
Ad-grade output generally, per invideo10 to 402.5% to 10%
Note

The first three bands are Poppify's; the last is invideo's. Practitioners describe experienced operators reaching a deliverable in a fraction of the generations a beginner burns, though I have no sourced figure for the gap.

Published runs are self-selected, so read them as existence proofs. invideo's documented UGC batch generated 108 images and 103 videos to land 49 used clips, roughly half the video generations, and the same page notes rejection hitting about 85% on harder ads, while the Kalshi NBA Finals spot is reported as 300 to 400 generations for about 15 usable clips, self-reported by the creator. Nobody publishes the run where they burned 400 generations and shipped nothing, and yield varies more by model than the price table suggests, the axis I compare in the model comparison for ad work.

Where the money goes

One itemised stack for a finished 60-second ad, vendor-adjacent but internally consistent:

LineCostShare
Editing labour, 2h at $50/h$100.0044%
Video compute, with 3x discard built in$72.0031%
Revision contingency, 20%$38.3017%
Ops overhead$15.007%
Music licensing$2.50
Voiceover$1.20
Image assets$0.80
Total$229.80

Writing and directing appear at zero, the flaw in every vendor cost comparison including this one. The Kalshi spot cost about $2,000 and took three days, and the man who made it was earning a living on live-action contracts before he moved to AI work.

Cost per winner is the only unit that decides anything

Nobody joins these datasets, so take a Small-tier account shipping about 35 finished ads a month, roughly the tier's top quartile at 8.0 a week, where Motion counts 0.5 winners. Price those 35 three ways: invideo's batch figure of about $125 an ad, including roughly two hours of operator time, is $4,375; the $145 invideo reports for the localisation and character-swap case is $5,075; Icon's published human rate of $1,000 for six ads is about $5,833. All three land between roughly $8,750 and $11,700 per counted winner, so the methods cost about the same at that volume and the case for generation rests on turnaround, shot types you cannot film, and volume beyond what a human roster delivers, which suits a dropshipper cycling new products weekly better than a brand with one hero SKU.

Icon is no neutral benchmark either. Its comparison page reads: "Icon originally launched as The AI Admaker. Today, we make 6 Human UGC for $1000 (no AI / 100% real)", backed by 1,288+ vetted creators and 200+ in-house editors. The best-funded company in AI UGC walking back to human filming is revealed preference, worth more than any survey here, and my reading is that the pivot buys reliability at low volume, where the arithmetic says human UGC wins anyway.

What none of this measures

Nothing cited above measures profit. Clicks, spend concentration, a new-adopter targeting number and the 0% to 16.3% sales-lift band are as close as it gets, since no public dataset on AI-generated creative carries an incrementality test, a geo-lift, a holdout or a margin, and every business case including mine assumes hit rate holds when you change production method.

Which ad formats earn the spend?

Motion's format table, again Motion's own data measuring spend concentration with ROAS out of scope:

FormatHit rateShare of creativesShare of spendSpend-use ratio
Unboxing9.8%2.1%2.8%1.3
Offer-First Banner8.6%21.9%29.3%1.3
Demo8.1%12.6%12.9%1.0
Testimonial6.5%13.3%13.3%1.0
Celebrity5.9%0.8%1.8%2.1

The window spans Black Friday and the post-holiday reset, which Motion flags as one of the most competitive promotion cycles of the year, so peak gifting probably flatters unboxing's top hit rate as much as the offer-led formats, and its 2% volume share may just mean you need a box worth opening. Read that row as a prompt to test unboxing in a giftable category.

Where the savings are real

The savings sit in variant production at a fixed concept: localise a proven winner into twenty languages, swap the product inside a working structure, test five hooks against an identical body, re-cut the same footage for a different placement. The marginal clip approaches zero, and every credible first-hand account converges here and on imagery you cannot shoot at all.

Four things AI video still cannot do:

  • Brand-consistent character work. Silverside's Svedka Super Bowl spot took about four months to reconstruct the Fembot character and train the models, per TechCrunch, longer than a conventional shoot and the most damaging single fact for the speed thesis.
  • On-screen text. Signage, prices, legal lines and small labels are unreliable across every model, so composite brand assets in post and brief the generation to end on a clean plate.
  • High-AOV considered purchases and emotional dramatic work, where the craft gap Ipsos measured is the one you cannot brief past.
  • Anything needing a specific real person, which is a legal problem before it is a technical one.

Dollar Shave Club is the useful model, producing about 90% of its advertising in-house with AI tooling while staying selective: its military campaign uses real footage because authenticity is the message, and a campaign that would have been impractical to shoot is fully generated. Shoot it where the realness is the meaning.

The prompt grammar that survives contact with a client

This comes from shipped pipelines rather than published documentation, so treat it as one operator's method, from someone selling no tools, courses or generation credits.

The trick that changed my output most costs nothing. Instead of generating eight clips and fighting drift between them, build one wide 21:9 sheet holding eight vertical 9:16 panels, feed it as the reference, and the model renders one continuous clip with eight internal hard cuts, panel K becoming cut K. All eight panels render in one image pass against the same references, so wardrobe, face, lighting and product hold by construction, and each reference gets one job with the rest forbidden: `@Image1 controls only the product geometry. Do not copy @Image1's studio background.`

Order the prompt as style and mood, a one-sentence summary, the cut-by-cut description with timecodes, the static setting, audio, then the negative suffix, and write the literal string `Hard cut to.` between consecutive cuts, seven markers for eight cuts, or they collapse into morphing motion. Word density is the single most useful number in the pipeline:

Clip lengthSpoken words
Up to 10 seconds12 to 20
11 to 12 seconds20 to 28
13 to 15 seconds28 to 35

A 30-second ad is therefore about 60 to 70 spoken words, and every writer I hand this to writes triple that on the first pass, when visual beats cost zero words and every plot event that can be shown should be. Three lines carry most of the realism: cuts only snap when adjacent beats differ on point of view, framing distance and physical action at once, more than two simultaneous hand roles writes a phantom third hand, and one clause fixes the most obvious tell in AI UGC, her mouth moves only during her own lines. The full syntax and the failure catalogue are in the Seedance 2.5 prompting walkthrough.

Frozen-frame QA before anything ships

Freeze evenly spaced frames, every product close-up, and two or three mid-word frames:

  1. Exactly one hero product, no clones
  2. At most two hands per person, including at frame edges
  3. Absent features stay absent, and cap, button and prop states stay consistent across cuts
  4. Labels are not gibberish, not mirrored, and not a different real brand
  5. Product scale matches the holding hand
  6. No doubled lip edges, no face drift, no baked text or subtitles

Item four is a legal check wearing a polish check's clothes.

Do you have to disclose that an ad is AI-generated?

Reviewed 10 September 2026, and this area moves fast enough to re-check before relying on it. For an ordinary US commercial short-form video ad no platform imposes an advertiser disclosure duty, though state law reaches this creative directly and the FTC's endorsement rules never needed to mention AI to apply.

WhereWhat is required
Meta, ordinary commercial adNo advertiser duty. Meta labels ad images itself when created or significantly edited with its own or third-party AI tools, excluding resizing and colour correction
Meta, social issue, election or political adsAdvertisers "are already required to disclose if the image, video or audio are created or edited with AI"
Google Ads, election adsDisclosure required, and Google auto-generates it on mobile Feeds, Shorts and in-stream while you write it yourself elsewhere. Cropping, colour correction and defect removal are exempt
YouTubeDisclose realistic altered or synthetic content: real people appearing to say or do things they did not, altered footage of real events, or realistic scenes that never occurred. Beauty filters and cloning your own voice are exempt
TikTokCreators must label realistic AI-generated content, and TikTok reads C2PA Content Credentials and auto-labels regardless of what you declared
New YorkGBL §396-b, in force 9 June 2026, requires conspicuous disclosure that a "synthetic performer" appears in an ad where the advertiser has actual knowledge. $1,000 first violation, $5,000 subsequent
CaliforniaSB 942 as amended by AB 853, operative 2 August 2026 at $5,000 per violation per day. Duties fall on large AI providers and, from 1 January 2027, on large online platforms, which must not strip provenance. No advertiser duty, though your output arrives carrying provenance you cannot strip
EU, any audienceArticle 50(4) requires deepfake disclosure by the deployer, which is you. It applied from 2 August 2026, enforcement began the same day, and penalties sit in the €15M or 3% of turnover tier. Active amendment and delay proposals make the current position worth confirming
ChinaVisible and metadata labels required since 1 September 2025 under the CAC labelling measures
Anywhere, synthetic testimonialThe FTC's rule on reviews and testimonials expressly prohibits misrepresenting that a reviewer or testimonialist exists

The Endorsement Guides are the piece most AI UGC pipelines miss. An endorsement is any message consumers are likely to believe reflects the opinions or experience of a party other than the advertiser, and an endorser "could be or appear to be" an individual, so the test turns on consumer perception rather than on whether the endorser exists. Plenty of blogs overstate this, since the regulation carries no example addressing AI-created endorsers and the FTC's own endorsement FAQ does not mention them, but it applies to a synthetic spokesperson anyway: material connections must be disclosed, a visual endorsement needs a visual disclosure, and exceptional results need typical-results context. The agency set aside its Rytr order in December 2025 as unduly burdening AI innovation, which pushes liability onto whoever publishes the claim.

Different obligations, different remedies, and brands get sued confusing the two. Tennessee's ELVIS Act defines a protected voice as a sound readily identifiable and attributable to an individual whether it contains the actual voice or a simulation, and attaches liability for advertising use without consent. The live test case is John R. Cash Revocable Trust v. The Coca-Cola Company, filed November 2025 over a national ad that allegedly used a soundalike singer and still at the pleading stage as at 10 September 2026. The alleged soundalike was a human imitator, so the statute attaches liability to the result without caring how you produced it.

In California, Civil Code §3344 covers knowing use of a real person's name, voice, photograph or likeness in advertising without consent, at the greater of $750 or actual damages plus attributable profits, with attorney's fees to the prevailing party, and it has no soundalike subsection, so imitations run through common-law misappropriation. Consent can evaporate too: under Labor Code §927 a digital-replica provision is unenforceable where it lacks "a reasonably specific description of the intended uses" and the individual had neither counsel nor a union agreement covering replicas, which perversely makes a broad grant worse than a narrow one.

Trademark and music, the two fastest ways to get an ad pulled

Trademark is the fast lane, because copyright cases grind for years while Cameo obtained a preliminary injunction against OpenAI over the naming of a generative video feature within about four months of filing, provisional relief rather than a merits holding. In the UK, Getty abandoned its training and output claims mid-trial and lost on secondary copyright infringement, but won in part on trade marks because Stable Diffusion generated images carrying Getty watermarks, a first-instance judgment subject to appeal. The live risk is a brand indicator leaking into your deliverable, a competitor's logo on a background object or a stock watermark you did not notice, which is what item four of the QA checks for, and Google's indemnity for generated output carves out trademark claims from use in trade or commerce, which advertising is by definition.

Music is the weakest link, because no major generator indemnifies your output. Suno's terms assign whatever rights Suno holds to paid subscribers and then state that Suno makes no representation that any copyright vests in it at all, and the US Copyright Office's position is that purely AI-generated material is not protected, so a generated track buys no exclusivity if a competitor uses an identical one. Brief music by mood, tempo, instrumentation and era, because a brief naming a reference artist breaches vendor terms, is discoverable, and supplies the knowledge element for a likeness claim.

What does an AI ad agency charge, and why?

PJ Accetturo, who made the Kalshi spot, charges businesses five-figure fees on production costs he says cap at around $2,000 per video, as he told Business Insider in an interview AOL carries, and six-figure per-spot numbers circulate for AI studios that I could not verify. That spread is the market rate for creative direction, reliability, legal cover and someone to carry the blame, and it reads to me as a talent-scarcity rent that compresses as directing supply catches up. Across published lists of the highest-volume DTC performance creative shops, none is described as running AI video generation as its primary engine, and TubeScience, reportedly the largest at roughly $2bn in managed spend, runs on statistical testing infrastructure.

Do people hate AI ads?

Four studies point the same way, all four discountable for method: the IAB's AI Ad Gap survey of 505 US Gen Z and Millennial consumers and 104 ad executives found 82% of executives believing consumers feel positive about AI ads against 45% of consumers who do, with Gen Z negative sentiment nearly doubling from 21% to 39% between the 2024 and 2026 waves; Gartner reported in March 2026 that half of US consumers would prefer brands avoid generative AI in customer-facing content, from 1,539 consumers fielded in October 2025; experiments at the Nuremberg Institute for Market Decisions, n = 1,000 in each of the US, UK and Germany, found the same ad rated more negatively when labelled AI, especially on emotional dimensions; and an NYU Stern finding in the IAB's disclosure framework puts the click-through cost of an AI label at 31.5%, via PPC Land, making disclosure a commercial decision as much as a compliance one.

The best-documented backlash case has an uncomfortable postscript, since Coca-Cola's AI Christmas work drew two consecutive years of public criticism and System1 scored it 5.9 stars, the top of their scale. For all the noise, I cannot find a documented sales decline, share-price move or measured brand-equity loss attributable to any AI ad backlash, which is either evidence the risk is overrated or evidence nobody has measured it properly.

A 30-day plan to a first AI ad worth scaling

  1. Days 1 to 3, fix the offer, because if it would not convert on a plain landing page with no ad at all you are about to spend a month making a bad offer more visible.
  2. Days 4 to 6, write concepts rather than variations, 12 to 20 spoken words per ten-second clip.
  3. Days 7 to 10, build the references once, multi-angle product shots and one canonical product description reused verbatim everywhere downstream, and sort the rights first: no real person's photograph as a character reference without a written digital-replica licence describing the intended uses specifically; check category, term and territory limits on any licensed AI actor before a localisation run, because a twenty-language campaign routinely exceeds them; and treat a fully synthetic face as safer on publicity risk while still carrying New York's label duty and the FTC endorsement rules.
  4. Days 11 to 16, generate against the de-slop pass and run the frozen-frame QA on every clip before it enters the edit.
  5. Days 17 to 18, size the test against your budget. Twelve ads in 3 to 5-variant ad sets gives two to four ad sets, and at 2 to 3 conversions per ad set per day over seven days that is 3 to 5 conversions per variant against a 30-to-50 threshold. Reaching 30 across twelve variants needs roughly 360 conversions, $18,000 at a $50 CPA, while the cycle as written costs about $3,500, so below roughly $20K a month ship 3 to 4 concepts instead.
  6. Days 19 to 26, run seven days without touching it. Hook and hold are impression-based and reach usable volume in days, so decide on those and treat CPA as directional until a variant has 30 to 50 conversions, which at small budgets takes weeks: the threshold and a 7-day read are incompatible on a small budget, and the honest resolution is fewer concepts. Never compare CPA across variants that received unequal budget.
  7. Days 27 to 30, graduate winners into existing scaling ad sets, and segment by day to check you are not scaling a novelty effect.

Motion counts 0.2 winners a month for the average account under $50K, so ten finished ads produces zero or one, and since nothing in the evidence tells you whether that winner is profitable, treat profit as the thing you are aiming at while the measurement available to you stops at clicks and spend.

Frequently asked questions

Do AI-generated ads perform better than human-made ads?

About the same in matched field data covering 369 million impressions, with no significant difference once campaign controls go in. Unadjusted, AI images that did not read as AI returned 0.79% click-through against 0.67% for human images that did not, while anything reading as AI fell to 0.62% or 0.55%.

Why do my AI ads look fake?

Most likely over-saturation, waxy skin, shallow depth of field and text baked into the frame. The research names saturation and displayed text as cues that signal AI, while high aesthetic scores and medium-to-large faces predict lower perceived artificiality. The depth-of-field and lighting advice is my own inference.

Do I have to disclose that an ad is AI-generated?

No US platform requires it for an ordinary commercial ad, but New York's GBL §396-b has required a conspicuous synthetic-performer disclosure since 9 June 2026, and disclosure is required for political and election ads on Meta and Google, realistic synthetic depictions on YouTube and TikTok, deepfakes shown to EU audiences, and any synthetic testimonial.

Does Meta penalise AI-generated ads?

No delivery or ranking penalty is documented anywhere I can find, and Meta labels AI-created or significantly AI-edited ad images itself while promoting AI creative through Advantage+. The cost of a label is reader response, not distribution, and TikTok and YouTube both say labelling does not affect reach.

Can I get fined for using AI-generated ads?

Not in the US for using AI as such. New York's synthetic-performer rule carries $1,000 and $5,000 penalties for a missing label, and EU Article 50(4) deepfake disclosure has applied since 2 August 2026 with penalties in the €15M or 3% of turnover tier. Realistic US exposure is the FTC endorsement and testimonial rules and likeness claims.

Is AI UGC better than hiring real creators?

At low volume, no. One vendor sells human-filmed UGC at about $167 an ad in a six-ad monthly package, beating most AI UGC subscriptions on price, and human filming does not dodge the aesthetic penalty either, since roughly a quarter of genuinely human images in the field data still read as AI. AI wins where volume is the constraint.

Sources

Every figure above traces to one of these. Where a source sells the thing it measured, we say so.

  1. 01AI in Disguise (working paper) · Exner, Hartmann, Netzer & Zhang, January 2025Not peer reviewed as of September 2026, and perceived artificiality was never randomly assigned.
  2. 02Cracking AI content creation · Harvard Business School AI Institute
  3. 03AI ads are indistinguishable from human ones, but they don't perform as well · Phys.org, on the Ipsos and Syracuse Newhouse study, May 2026Ipsos sells ad testing, and the study scored predicted effect on a survey panel rather than delivered clicks.
  4. 04Research: AI-generated ads perform worse than human-made ones, even when customers can't tell them apart · Harvard Business Review, September 2026
  5. 05Generative AI and personalised video advertising · Kapoor & Kumar, Marketing ScienceThe generative lift is confounded with personalisation.
  6. 06Field RCTs on generative AI in retail workflows, MSI Working Paper 26-112 · Fang et al., Marketing Science Institute, 2026
  7. 07AI-generated ads vs human ads: performance data · LapisLapis sells AI ad generation, and four of its five cited sources are vendors. Its own caveat is that the figures come from live campaigns rather than controlled A/B tests.
  8. 08Meta Advantage+ creative · Meta, Checked 10 September 2026First-party page selling the product it measures, with no methodology, sample or period attached, edited without version history.
  9. 09Andromeda: Advantage automation and next-gen personalized ads retrieval · Meta Engineering, December 2024The 22% ROAS figure is a new-adopter number about targeting, and the 7% image-generation figure is Meta's own estimate.
  10. 10Marketers vastly understate the sales effect of creative and significantly overestimate the impact of targeting · Westwood One, summarising NCSolutions and Nielsen, April 2024
  11. 11When it comes to advertising effectiveness, what is key? · Nielsen, 2017
  12. 12NIQ research uncovers hidden consumer attitudes toward AI-generated ads · NIQ, December 2024A press release from a company selling consumer research, with no per-format numbers and no stimuli published. Every current video model postdates it.
  13. 13Moving to a positive attention economy with Attention-Adjusted Net Reach · Nelson-Field / Amplified Intelligence, via WARCAmplified Intelligence sells attention measurement, and the r = .82 correlation is published without a sample size.
  14. 14Hacking the attention economy: VCCP Media and Dr Karen Nelson-Field reveal the 1.5-second formula · VCCP Media, May 2025VCCP sells media planning, and no auditable method is published.
  15. 15Karen Nelson-Field on hacking the attention economy · Creative Salon
  16. 16Introducing the attention economy: attention research based on eye-tracking data of ad viewing · Dentsu JapanDentsu sells media planning, and the recall curve is published without an auditable method.
  17. 17ViASNet: a video-ad saliency model · Ye & Wedel, arXiv 2605.29302Stimuli are Dutch television commercials, so the numbers may not travel to vertical social video.
  18. 18Hook rate and hold rate benchmarks · Ad LibraryPractitioner aggregate, not an official platform benchmark.
  19. 19Hook rate benchmarks · Sepia LabThe widely repeated 30.7% TikTok average traces back to eleven accounts.
  20. 20Video play reporting metrics · TikTok Ads
  21. 21Creative best practices · TikTok AdsPlatform guidance from a company selling the inventory it advises on.
  22. 22Creative Benchmarks 2026: methodology · Motion, September 2025 to January 2026Motion sells creative analytics whose value proposition is that testing volume matters, and the data is not independently auditable.
  23. 23Creative Benchmarks 2026: testing volume by tier · Motion, September 2025 to January 2026Vendor-published, measures spend concentration rather than business outcomes, and no refresh has been published since.
  24. 24Creative Benchmarks 2026: top-quartile accounts · Motion, September 2025 to January 2026Vendor-published. Its counted winners disagree fivefold with the hit rates in the same report.
  25. 25Creative Benchmarks 2026: top visual formats · Motion, September 2025 to January 2026Vendor-published, ROAS out of scope, and the window spans Black Friday.
  26. 26Pitfalls of long-term online controlled experiments · Dmitriev et al., KDD 2017, 2017
  27. 27How to test ad creatives on Meta · Rocketship HQPractitioner consensus with no published methodology behind it.
  28. 28Meta creative testing campaign structure · Rocketship HQPractitioner consensus with no published methodology behind it.
  29. 29TikTok Next 2026 trend report: performance marketer playbook · SegwiseA vendor relaying another vendor's analysis, two vendors deep with no method published.
  30. 30The silent world of Facebook video · Digiday, 2016Source of the 85%-without-sound claim: self-reported numbers from three publishers, never Meta research, and it predates Reels.
  31. 31Kantar report: how brands are making noise with sound on TikTok · TikTok for Business, with KantarAttitudinal self-report, published by the platform selling the inventory.
  32. 32Ads guide: Facebook feed video · Meta
  33. 33Seedance 2.5 pricing survey · Cellcog, Verified against provider pages 22 August 2026
  34. 34AI video generator pricing · DIY AIIndependent cross-check that lands on the same ordering at roughly half the absolute figures.
  35. 35What an AI UGC ad actually costs · invideoinvideo sells generation credits and supplies both the yield figures and the per-ad cost.
  36. 36What percentage of AI-generated video clips are actually usable · invideoinvideo sells generation credits, and published runs are self-selected.
  37. 37The real cost of AI video generation · PoppifyVendor-published yield bands.
  38. 38AI video career and cost guide 2026 · AI Video Bootcamp, 2026Vendor-adjacent cost stack that prices writing and directing at zero.
  39. 39The chaotic Kalshi ad during the NBA Finals · YahooGeneration counts are self-reported by the creator.
  40. 40The filmmaker behind the AI-generated Kalshi ad on his fees and production costs · Business Insider, via AOLSelf-reported pricing from someone selling the service.
  41. 41SynthID · Google DeepMind
  42. 42Giving creators more control with Dreamina Seedance 2.5 and Dola Seedream 5.0 Pro · CapCut newsroomFirst-party announcement from the company shipping the watermarks and labels it describes.
  43. 43Arcads vs Creatify vs Icon · IconIcon sells human UGC and supplies the comparison price it is measured against.
  44. 44Top DTC performance creative agencies in 2026 · Yall
  45. 45Dollar Shave Club's bet that AI makes agencies optional, not obsolete · Digiday
  46. 46Super Bowl LX AI ads: Svedka, Anthropic and the brands · TechCrunch, February 2026
  47. 47The AI Ad Gap Widens · IABSurvey of 505 consumers and 104 ad executives, published by the trade body for the industry it surveys.
  48. 48Gartner marketing survey finds 50% of consumers prefer brands that avoid using GenAI in consumer-facing content · Gartner, March 20261,539 consumers, fielded October 2025, stated preference rather than observed behaviour.
  49. 49Transparency Without Trust · Nuremberg Institute for Market Decisionsn = 1,000 in each of the US, UK and Germany.
  50. 50AI ad labels cut click-through 31.5%: IAB framework cites NYU study · PPC LandSecondary reporting of an NYU Stern finding carried inside an IAB disclosure framework.
  51. 51AI or no AI, Coke gets the Christmas love · System1System1 sells ad testing, and the 5.9-star score is its own proprietary scale.
  52. 52AI labelling in ads · Meta Help Centre
  53. 53Political content policy · Google Ads Help
  54. 54Disclosing altered or synthetic content · YouTube Help
  55. 55Partnering with our industry to advance AI transparency and literacy · TikTok Newsroom
  56. 56General Business Law §396-b, synthetic performer disclosure · New York State Senate, In force 9 June 2026
  57. 57AB 853, the California AI Transparency Act · California Legislative Information, Operative 2 August 2026
  58. 58Civil Code §3344 · California Legislative Information
  59. 59Labor Code §927, digital replicas · California Legislative Information
  60. 60EU AI Act, Article 50 · artificialintelligenceact.euActive amendment and delay proposals make the current position worth confirming.
  61. 61EU AI Act, Article 99 penalties · artificialintelligenceact.eu
  62. 62Commission starts enforcing AI Act rules and new transparency requirements · European Commission, 2 August 2026
  63. 63Measures for labelling AI-generated synthetic content · Cyberspace Administration of China, In force 1 September 2025
  64. 64Federal Trade Commission announces final rule banning fake reviews and testimonials · Federal Trade Commission, August 2024
  65. 65Endorsement Guides, 16 CFR Part 255 · Cornell Legal Information Institute
  66. 66The FTC's Endorsement Guides: what people are asking · Federal Trade CommissionCarries no example addressing AI-created endorsers.
  67. 67FTC reopens and sets aside the Rytr final order in response to the AI Action Plan · Federal Trade Commission, December 2025
  68. 68ELVIS Act, HB 2091 · Tennessee General Assembly
  69. 69John R. Cash Revocable Trust v. The Coca-Cola Company, docket · CourtListener, Filed November 2025Still at the pleading stage as at 10 September 2026.
  70. 70Baron App (Cameo) v. OpenAI, docket · CourtListenerPreliminary injunction is provisional relief, not a merits holding.
  71. 71Getty Images v Stability AI [2025] EWHC 2863 (Ch) · The National Archives, Find Case LawFirst-instance judgment subject to appeal.
  72. 72Generative AI indemnified services · Google CloudThe indemnity carves out trademark claims arising from use in trade or commerce.
  73. 73Terms of service · SunoSuno assigns whatever rights it holds and then states it makes no representation that any copyright vests in it.
  74. 74Copyright and artificial intelligence · US Copyright Office