Running 50+ AdMob Mediation Groups Alone: Four Tasks I Handed to an Agent, Four Decisions I Kept
Two wallpaper apps on iOS and Android put me past 50 AdMob mediation groups. Here are the four tasks I handed to an agent, the four decisions I kept, and the sample-size hole I later found in the evaluation code I had published.
Across my two wallpaper apps on iOS and Android, I now run more than 50 AdMob mediation groups. BW iOS has 28, BW Android 14, UKI iOS 16, UKI Android 6. One mediation per app multiplied into this many after App Tracking Transparency (ATT) ON/OFF splits and regional splits piled on.
The time I spent each morning opening AdMob reports and nudging floor values by hand kept creeping up. One morning my coffee went cold while I scrolled the same screen, and it struck me that the problem was not the volume of work — it was that I had never written down the procedure. So over the following month I drew a line between what an agent can run and what I refuse to delegate.
If you are an indie dev with 10+ mediation groups, the tweaks are eating non-trivial time, and you'd like to put an agent in the loop without losing revenue, this is written for you.
Why the group count grew this far
Before the agent discussion, a quick map of how I ended up at 50+. Modern AdMob mediation practice splits groups along three axes: ATT ON/OFF × top countries × format (INT/BNR/RWD/RWI/AOA). For BW iOS interstitials, the split looks like:
BW-INT-iOS-JP-ATT-ON / OFF
BW-INT-iOS-TW-ATT-ON / OFF
BW-INT-iOS-US-ATT-ON / OFF
BW-INT-iOS-IT-ATT-ON / OFF
BW-INT-iOS-DE-ATT-ON / OFF
BW-INT-iOS-HK-ATT-ON / OFF
BW-INT-iOS-ROW-ATT-ON / OFF
7 regions × 2 ATT values = 14 groups for one format on one platform. Add BNR / RWD / RWI / AOA and you cross 30 per platform, 60+ across both OS. Realized eCPM ranges from $1 to $34 across groups, each with its own optimal floor.
A "floor" here is the minimum eCPM across the full bidding + waterfall stack. Raise the floor and unit price climbs but match rate falls. Lower it and the inverse. My rule of thumb: floor = realized eCPM × 50–60%; revisit when the floor-to-eCPM ratio crosses 65%.
Tasks I now hand off — the four I let an agent run
After a month of experiments, the four tasks I trust an agent with are these.
The first is match-rate monitoring. The agent opens the AdMob report at a 30-day window and aggregates ad requests, impressions, match rate, and estimated revenue per group. I issue this style of instruction:
On the AdMob report:1. Window: last 30 days2. Dimensions: app + mediation group3. Metrics: ad requests, impressions, match rate, estimated revenue4. Filter: groups with matchrate < 80% only5. Output: group name / current floor / realized eCPM / floor-to-eCPM ratio / ad request count / recommended action
The agent pulls the data via API or CSV export and returns a table. I read the table and decide "keep JP, adjust ROW." Aggregation dropped from 30 minutes to 3 minutes.
The second is floor-to-eCPM ratio checks. Trivial math, but the place where humans make small errors at scale:
def evaluate_floor(group_name, current_floor, ecpm_30d, matchrate): """Score the relationship between floor and realized eCPM.""" ratio = current_floor / ecpm_30d if matchrate < 0.80: if ratio > 0.65: return f"{group_name}: ratio={ratio:.2%} — candidate to lower (matchrate {matchrate:.0%})" return f"{group_name}: ratio={ratio:.2%} — likely demand shortage (floor not the cause)" if ratio > 0.65: return f"{group_name}: ratio={ratio:.2%} — ratio high but matchrate OK, hold" return f"{group_name}: healthy"# Expected output:# BW-INT-AND-JP: ratio high but matchrate OK, hold# BW-INT-AND-DE: ratio=88.96% — candidate to lower (matchrate 56%)
There are two holes in this function that I only found later. I fix both in a section further down, so read on with this version as it stands.
The third is operations-log updates. Every floor change gets appended to admob-mediation-unified.md with date, group ID, before → after, and reason. I describe the change out loud; the agent files the entry. The asset is being able to ask future-me, "why did I lower the JP floor from $15 to $12 on 2026-05-17?" and getting an answer in seconds.
The fourth is applying the "half rule" for AppLovin MAX. When AppLovin MAX joins a waterfall already running Unity Ads, I set AppLovin's manual eCPM to half of Unity's. Unity at $8.50 → AppLovin $4.25; Unity at $5.00 → $2.50. Computing 20 groups by hand will produce at least one error.
Read the Unity floor column below. Compute the AppLovinMax waterfallfloor at half. Output a paste-ready number-only list for AdMob's editscreen. For BNR groups, ignore the Unity value and output a fixed $0.15.| Group ID | Unity floor (waterfall) || BW-INT-AND-JP (8292611016) | $8.50 || BW-INT-AND-KO (5982272500) | $5.00 || ... |
The output goes directly into AdMob's edit screen. Ten groups in under a minute, and the error rate of manual arithmetic disappears.
✦
Thank you for reading this far.
Continue Reading
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦You can add a hold rule to your own evaluation code so small-region groups are not changed on noisy numbers
✦You can write the rollback condition into your ledger first, so a seasonal swing is not mistaken for your own win
✦You can sort your own mediation chores into what an agent may run and what you must keep, before trying to shrink the monthly review
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
Rereading evaluate_floor() some time after writing it, my hand stopped. It compares match rate to a threshold as a point estimate.
Match rate is a proportion — how many of your ad requests got filled — so the fewer the requests, the wilder the number. I had been cutting a small regional group with 172 requests in 30 days and a group with 2,000 requests along the very same 80% line.
Here is how far apart those two cases actually are, using a 95% Wilson interval. For an observed match rate of 75%:
Ad requests
Observed
95% interval
Half-width
Can you claim "below 80%"?
30
75.0%
55.6% – 85.8%
15.1pt
No
100
75.0%
65.7% – 82.5%
8.4pt
No
200
75.0%
68.6% – 80.5%
6.0pt
No
300
75.0%
69.8% – 79.6%
4.9pt
Yes
1,000
75.0%
72.2% – 77.6%
2.7pt
Yes
5,000
75.0%
73.8% – 76.2%
1.2pt
Yes
At an observed 75%, you need 230 requests before you can say "this is under 80%" with 95% confidence. At an observed 78% — closer to the target — you have to wait for 1,498 requests before the interval's upper bound clears 80%. At an observed 60%, fourteen requests are enough. The bigger the gap, the smaller the sample you need, which sounds obvious until it turns into weeks of waiting in your monthly loop.
The second hole was simpler: ecpm_30d at zero or missing makes the division blow up. That happens for real on BNR groups that only just started serving.
Here is the corrected version.
from math import sqrtZ95 = 1.96def wilson(matched, requested, z=Z95): """Return the (lower, upper) confidence bounds for a match rate.""" if requested <= 0: return (0.0, 1.0) p = matched / requested d = 1 + z * z / requested center = (p + z * z / (2 * requested)) / d margin = z * sqrt(p * (1 - p) / requested + z * z / (4 * requested * requested)) / d return (center - margin, center + margin)def evaluate_floor(group, floor, ecpm_30d, matched, requested, target=0.80, min_requests=200): if not ecpm_30d: return f"{group}: eCPM is zero or missing — cannot evaluate" if requested < min_requests: return f"{group}: requests={requested} — held, sample too small" ratio = floor / ecpm_30d lo, hi = wilson(matched, requested) if hi < target: if ratio > 0.65: return f"{group}: ratio={ratio:.0%} / CI={lo:.0%}-{hi:.0%} — candidate to lower" return f"{group}: ratio={ratio:.0%} / CI={lo:.0%}-{hi:.0%} — likely demand shortage" if lo >= target: if ratio > 0.65: return f"{group}: ratio={ratio:.0%} — ratio high but matchrate at target, hold" return f"{group}: healthy" return f"{group}: CI={lo:.0%}-{hi:.0%} straddles 80% — held"
Running five groups through it locally:
BW-INT-AND-DE: ratio=178% / CI=54%-58% — candidate to lowerBW-INT-AND-JP: ratio=98% / CI=43%-47% — candidate to lowerUKI-RWD-iOS-HK-ATT-OFF: requests=172 — held, sample too smallUKI-BNR-iOS-IT-ATT-ON: eCPM is zero or missing — cannot evaluateBW-INT-AND-TW: CI=71%-81% straddles 80% — held
The last two lines are the point. UKI-RWD-iOS-HK has only 172 requests, so it falls into the held bucket. BW-INT-AND-TW matched 214 of 280 requests — an observed 76% that the old code would have listed as "candidate to lower" without hesitation — yet its interval runs 71%–81% and crosses the line. Drop the floor in that state and, if match rate recovers next month, I may well credit the change for something the sample never established.
Act on whether the threshold was crossed provably, not on whether it was crossed. Since adding that rule, the number of groups I touch each month has gone down. What changed most, I think, is that "leave it alone" now shows up in the output as a real answer.
What I refuse to hand over — four domains I keep
The other side of the line. These cases satisfy at least one of "directly affects revenue," "data alone is insufficient," or "embedded business logic."
The first is setting absolute floor values. Even if the agent proposes "lower JP from $15 to $10.50," the final approval is mine. Floor decisions track not only realized eCPM but also seasonal ad demand (year-end retail, spring fresh starts) and the wave of specific networks (Meta, AdMob Network). Data alone won't surface those.
The second is wholesale stopping or restarting a network. On 2026-05-26 I paused AppLovin Bidding and Waterfall across all of my apps on INT/RWD/RWI. The trigger was a CTR 13% × RPC $0.0092 × eCPM $1.21 anomaly on BW Android, plus a long crash trend in App Store Connect showing a clear spike in the period when AppLovin was first integrated. That kind of decision pulls from data, from watching test ads, and from memory an agent doesn't hold. About $730/month of revenue went away. That's a call I make and own.
The third is interpreting anomalies. When match rate moves materially week-over-week, or one group's eCPM doubles overnight, attributing it to "network bidding logic change," "my own release timing," or "season" is hard for an agent. In my workflow, I lay Crashlytics, Play Console, App Store Connect, and AdMob reports side by side and reason it out; the context only lives on the human side here.
The fourth is adding or removing an ad SDK. Whether to integrate Mintegral, submit InMobi for review, or raise Meta's priority touches not just revenue but crash risk, UX, and apk footprint. If the agent says "Meta has higher average eCPM than AdMob Network," I don't promote Meta on that basis alone. Knowing Meta's quirks (the Code 1002 too frequently frequency-cap behaviour) matters more than the average eCPM number.
How the handoff actually runs — a monthly review loop
I run a scheduled monthly review across all 50+ groups. The agent runs:
[Step 1] Data pull from AdMob report - Window: last 30 days - Ad requests / impressions / match rate / estimated revenue / realized eCPM for all 50+ groups - Persist results as structured data (JSON / CSV)[Step 2] Programmatic floor-to-eCPM evaluation - Call evaluate_floor() on every group - Bucket into "candidate to lower," "candidate to raise," "hold," "likely demand shortage," "held, sample too small"[Step 3] Month-over-month diff - Compare against last month's operations log - Highlight groups where eCPM moved more than ±20%[Step 4] Report generation - Top 5 revenue contributors - Bottom 5 match rates, with intervals - "Needs human judgment" list - Mechanically applicable change candidates
I read the report in the morning with coffee, approve the subset I'm willing to commit, and tell the agent to "execute." The agent either calls APIs or drives the AdMob UI through a browser tool, then appends entries to admob-mediation-unified.md and we're done.
Before the agent ran this loop, it was 3–4 hours of manual work per month. Now it is under 30 minutes. Those 3 hours go into app-side code maintenance and review replies — work agents can't do well for me yet.
A failure case I want to flag
Worth writing down, because the same trap might catch someone else.
The agent once proposed temporarily setting a group's floor to OFF to recover match rate. I approved it. The minimum eCPM disappeared across the entire bidding + waterfall stack, and overall eCPM caved. The AdMob spec is clear that Waterfall manual eCPM is not a substitute for the floor, but acting on an agent's proposal without re-reading the spec is easy to do.
After that I wrote a rules file (feedback_admob_ecpm_floor_warning.md) and started copy-pasting it into every agent instruction:
[Prohibited]- Never set an INT group's floor to OFF- Disabling the floor erases the minimum eCPM across bidding too- Waterfall manual eCPM is NOT a substitute for the floor
The safer pattern is to hand the agent both "what to do" and "what never to do," explicitly. Looking back, approving that change before any prohibition list existed was the actual mistake.
Recent tuning, by way of example
A few real adjustments, labelled by agent involvement:
2026-05-17 (Android BW-INT) — 30-day eCPMs landed at JP $15.23, KO $8.64, TW $8.87, DE $5.62. Match rates JP 45% / KO 47% / TW 67% / DE 56%, all under 80%. The agent listed them as "floor too high" candidates → I approved → floors changed JP $15→$12, KO $10→$7, TW $10→$8, DE $10→$5
2026-05-24 (BW/UKI Android, all 20 groups) — Applied the half rule for AppLovin MAX. Agent computed → I pasted into AdMob's edit screen for 20 groups
2026-05-26 (All iOS/Android AppLovin INT/RWD/RWI pause) — Decision driven by crash evidence and CTR anomaly. Pausing was a manual click; the agent organised applovin-stop-decision-2026-05-26.md
That last one is the right shape for large decisions: agent organises the evidence, human pulls the trigger. Aggregation and reasoning go to the machine; the actual button press stays human.
One caveat on the first row: the current evaluation code would not reach the same verdict for every group there. JP at 45% with plenty of requests still lands as "candidate to lower," but the thinner regions fall into the held bucket. Back then I couldn't tell the difference between numbers being present and numbers being sound.
Write the rollback criteria before you apply the change
Running the monthly review taught me that "when to roll back" is vaguer than "when to lower." If match rate rises the month after a floor cut, it is tempting to credit the cut — even though the cause may be seasonal.
So on the day a change goes in, I now write the rollback criteria into the ledger as well. The function below takes the requests and matches accumulated since the change and returns "not yet," "roll back," or "keep." It reuses wilson() from earlier.
from datetime import datedef review_change(group, matched_after, requested_after, expected=0.80, min_requests=200, opened=None, today=None, max_days=45): """Post-change review. Executes exactly the criteria written in the ledger.""" opened = opened or date.today() today = today or date.today() elapsed = (today - opened).days if requested_after < min_requests: if elapsed >= max_days: return f"{group}: still thin after {elapsed} days - record 'effect not measurable' and close" return f"{group}: requests={requested_after} - not judging yet" lo, hi = wilson(matched_after, requested_after) if hi < expected: return f"{group}: CI={lo:.0%}-{hi:.0%} - below expectation, candidate to restore the old floor" if lo >= expected: return f"{group}: CI={lo:.0%}-{hi:.0%} - as expected, keep" return f"{group}: CI={lo:.0%}-{hi:.0%} - interval straddles the target, wait another cycle"
Since adding this, the ledger entry has two extra fields after "before → after": the review date and the rollback condition.
Ledger field
What goes in it
Why it is there
Before → after
A concrete value, e.g. $12 → $10
So the way back is never in doubt
Expected metric
e.g. match rate 80%
So the bar does not move after the change
Review date
30–45 days after the change
So I do not touch it before the sample accumulates
Rollback condition
CI upper bound below the target
So I do not roll back on a hunch
If it never measures
Write "effect unknown" and close it
So small regions stop getting nudged
The last row does the most work. A region with few requests never narrows its interval, however long I wait. Writing "could not be measured" and closing the entry quietly removes it from the monthly review. Not touching something is also a decision worth recording.
Instead of making decisions faster, set things up so I do not need to hurry them. The same line helped when I was deciding what to hand over.
Starting "half-agented" mediation operations
For someone bringing an agent into AdMob mediation today, here's the first-week order:
Consolidate the existing mediation groups into a single operations log (admob-mediation-unified.md or equivalent)
Write the floor-to-eCPM evaluation function in Python, run it locally on your own data first
Put a minimum ad-request count (200–300 is a reasonable start) into that function from day one
Lock in one prompt for monthly match-rate monitoring, and schedule it
Write down a "prohibited" list and always include it in agent instructions
Run the first month in "proposals only, human executes" mode
Once a month runs cleanly, promote the low-risk tasks (operations log entries, half-rule arithmetic) to autonomous execution
Running AdMob on my own has taught me that monetization tuning matters more for reproducibility than for raw speed. A lucky eCPM bump this month means nothing if next month's path can't be reproduced. The point of bringing an agent in is to document your own judgment so that anyone, including future-you, could run the same loop and reach the same outcome. The time savings are a side effect.
When I repair photographs of old paper for the ukiyo-e wallpaper app, I decide how far to go before I touch anything. Start without deciding, and you look up to find you've added things the paper never had. Mediation operations have the same shape: the drawing — the operations log and the decision rules — comes first, and the agent's judgment stays inside it.
Start by adding a requests column to your own evaluation function and counting how many groups are sitting below the sample threshold. That's where I restarted. Thank you for staying with me through the interval math.
Share
Thank You for Reading
Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.