◉ANTIGRAVITY LABJP
Articles/AI Tools
⚙ AI Tools/2026-04-22Advanced

Running Multiple Gemma 4 LoRAs in Production — A Practical Guide to Merging and Dynamic Adapter Switching

You've trained several LoRAs on Gemma 4 — summarization, translation, code review. How do you serve them without tripling your GPU bill? A working notebook on pre-merge rank checks, Weighted and TIES merges, dynamic switching, and measuring switch churn, written with Antigravity alongside.

gemma-419lora4mergekitpeftmulti-taskantigravity461

✦ Premium Article

The morning after I finished training three LoRAs

One LoRA to summarize app reviews. One to translate store descriptions between Japanese and English. One for code review. As an indie developer, those are three chores that show up in my week over and over, so I trained a separate adapter for each on top of Gemma 4. Each one looked good on its own evaluation, and I went to bed pleased with the weekend.

The next morning I sketched out what it would cost to keep three separate servers running, and my stomach tightened a little. Review summaries cluster in the morning, translation happens right before a store update, code review happens at night. Paying for three GPUs around the clock for three tasks that rarely overlap simply doesn't fit a solo developer's budget.

This is the working notebook from that stretch. It covers two practical approaches — LoRA merging and dynamic adapter switching — and walks through the places I actually got stuck, with Antigravity's agent riding along: how Weighted and TIES differ, what to check before any merge, and the set_adapter() trap.

If you haven't trained a LoRA yet, start with LoRA / QLoRA fine-tuning for Gemma 4 first — this piece assumes you have at least two adapter checkpoints ready.

Three paths — and how I decide between them

When you have multiple LoRAs trained for different tasks, your options collapse into three:

  • Strategy A: Merge them into a single model. You fold the adapter weights into one checkpoint and deploy it like an ordinary Gemma 4 model. The inference path stays simple.
  • Strategy B: Keep one base model and swap adapters per request. Gemma 4 stays loaded; only the adapter changes, driven by task routing. You save memory but pay a switching cost and inherit a concurrency problem.
  • Strategy C: Activate multiple adapters in parallel. Dedicated servers like S-LoRA or Punica apply several LoRAs within one batch. Maximum flexibility, maximum operational weight.

My rule of thumb:

  • Two or three tasks with a similar character (JA→EN and EN→JA translation, summarization and extraction) → Strategy A
  • Four or more independent tasks (summarization, code review, SQL generation) → Strategy B
  • Tasks that blend inside a single response (summarize a technical doc and translate it in one go) → Strategy C

Four merge methods, and a simple order to try them in

  • Weighted (Linear): new_W = α * W_A + β * W_B. The oldest and simplest; good when the tasks barely interfere.
  • TIES (Trim, Elect Sign, Merge): drops small changes, resolves sign conflicts by majority vote, then merges. My first try whenever three or more adapters are involved.
  • DARE (Drop And REscale): randomly drops changes and rescales the rest. DARE-TIES becomes a candidate when you're merging many adapters.
  • SLERP: spherical interpolation between two models. Nice for blending tone; not usable for three or more.

Two adapters: Weighted or SLERP. Three or more: TIES. Many more: DARE-TIES. I run them in that order and keep whichever scores best on my own evaluation set. Lining a few candidates up side by side ends up faster than betting on one blindly.

✦

Thank you for reading this far.

Continue Reading

What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.

WHAT YOU'LL LEARN
✦If you've trained several LoRAs but weren't sure how to actually serve them, you'll walk away with three concrete production patterns you can choose between today
✦A preflight script that compares adapter_config.json files lets you catch rank and base-model mismatches before a merge ever runs
✦You'll run a dynamic-LoRA inference server and a switch-churn benchmark locally, so the strategy decision comes from your own traffic rather than someone else's numbers
Secure payment via Stripe · Cancel anytime
✦

Unlock This Article

Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.

or
Unlock all articles with Membership →
Share

Thank You for Reading

Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

Related Articles

⚙ AI Tools2026-04-21
Tuning Gemma 4 for Yourself — A Realistic LoRA / QLoRA Workflow on a Solo Developer's Budget
Full fine-tuning of Gemma 4 is out of reach for most individuals, but LoRA / QLoRA makes personalization realistic on a solo budget. This guide walks through data prep, training settings, evaluation, and wiring the result into an Antigravity workflow — from hard-earned practical experience.
⚙ AI Tools2026-06-12
Cutting Down 'Plausible but Wrong' RAG Answers — A Retrieval Evaluation Harness for Gemma 4 and Antigravity
Replace gut feeling with recall@5, MRR and faithfulness scores — a 30-question golden dataset and a small Python harness for evaluating a local Gemma 4 RAG stack.
⚙ AI Tools2026-10-01
You Don't Need to Read Commands: Three Places to Check When an Agent Works
If strings of terminal text make the approve button feel risky, here is a way to check just three things: a way back before you start, the first word of a command, and the diff afterward.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links