Running Multiple Gemma 4 LoRAs in Production — A Practical Guide to Merging and Dynamic Adapter Switching
You've trained several LoRAs on Gemma 4 — summarization, translation, code review. How do you serve them without tripling your GPU bill? A working notebook on pre-merge rank checks, Weighted and TIES merges, dynamic switching, and measuring switch churn, written with Antigravity alongside.
One LoRA to summarize app reviews. One to translate store descriptions between Japanese and English. One for code review. As an indie developer, those are three chores that show up in my week over and over, so I trained a separate adapter for each on top of Gemma 4. Each one looked good on its own evaluation, and I went to bed pleased with the weekend.
The next morning I sketched out what it would cost to keep three separate servers running, and my stomach tightened a little. Review summaries cluster in the morning, translation happens right before a store update, code review happens at night. Paying for three GPUs around the clock for three tasks that rarely overlap simply doesn't fit a solo developer's budget.
This is the working notebook from that stretch. It covers two practical approaches — LoRA merging and dynamic adapter switching — and walks through the places I actually got stuck, with Antigravity's agent riding along: how Weighted and TIES differ, what to check before any merge, and the set_adapter() trap.
If you haven't trained a LoRA yet, start with LoRA / QLoRA fine-tuning for Gemma 4 first — this piece assumes you have at least two adapter checkpoints ready.
Three paths — and how I decide between them
When you have multiple LoRAs trained for different tasks, your options collapse into three:
Strategy A: Merge them into a single model. You fold the adapter weights into one checkpoint and deploy it like an ordinary Gemma 4 model. The inference path stays simple.
Strategy B: Keep one base model and swap adapters per request. Gemma 4 stays loaded; only the adapter changes, driven by task routing. You save memory but pay a switching cost and inherit a concurrency problem.
Strategy C: Activate multiple adapters in parallel. Dedicated servers like S-LoRA or Punica apply several LoRAs within one batch. Maximum flexibility, maximum operational weight.
My rule of thumb:
Two or three tasks with a similar character (JA→EN and EN→JA translation, summarization and extraction) → Strategy A
Four or more independent tasks (summarization, code review, SQL generation) → Strategy B
Tasks that blend inside a single response (summarize a technical doc and translate it in one go) → Strategy C
Four merge methods, and a simple order to try them in
Weighted (Linear):new_W = α * W_A + β * W_B. The oldest and simplest; good when the tasks barely interfere.
TIES (Trim, Elect Sign, Merge): drops small changes, resolves sign conflicts by majority vote, then merges. My first try whenever three or more adapters are involved.
DARE (Drop And REscale): randomly drops changes and rescales the rest. DARE-TIES becomes a candidate when you're merging many adapters.
SLERP: spherical interpolation between two models. Nice for blending tone; not usable for three or more.
Two adapters: Weighted or SLERP. Three or more: TIES. Many more: DARE-TIES. I run them in that order and keep whichever scores best on my own evaluation set. Lining a few candidates up side by side ends up faster than betting on one blindly.
✦
Thank you for reading this far.
Continue Reading
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦If you've trained several LoRAs but weren't sure how to actually serve them, you'll walk away with three concrete production patterns you can choose between today
✦A preflight script that compares adapter_config.json files lets you catch rank and base-model mismatches before a merge ever runs
✦You'll run a dynamic-LoRA inference server and a switch-churn benchmark locally, so the strategy decision comes from your own traffic rather than someone else's numbers
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
Setting up with Antigravity — and starting with a pre-merge check
I tried this on a single-GPU CUDA box on Linux and on an Apple Silicon Mac, and the examples read the same on both.
When I asked Antigravity's agent to "install peft and write a minimal sanity check for merging and switching multiple Gemma 4 LoRAs," it scaffolded most of what follows. It also handed me my first mistake: the model ID it proposed was gemma-4-2b-it, which doesn't exist, and loading failed with a 404. The small Gemma 4 checkpoints are published with an "E" prefix, like google/gemma-4-E2B-it, and the model card's official loader is AutoModelForMultimodalLM, since these models also take images and audio.
# requirements.txt (verify versions against your own GPU and drivers)# torch# transformers # a release with Gemma 4 support# peft# accelerate# safetensors# httpx fastapi uvicorn # for the server and benchmark later onpython -m venv .venv && source .venv/bin/activatepip install -U -r requirements.txt# Log in to Hugging Face (needed to pull Gemma 4 weights)hf auth login# Sanity check: load Gemma 4 E2B and generate oncepython - <<'PY'from transformers import AutoProcessor, AutoModelForMultimodalLMMODEL_ID = "google/gemma-4-E2B-it"processor = AutoProcessor.from_pretrained(MODEL_ID)model = AutoModelForMultimodalLM.from_pretrained(MODEL_ID, dtype="auto", device_map="auto")messages = [{"role": "user", "content": "Summarize in one line: after the wallpaper app update, three reviews in a row said loading felt slow."}]inputs = processor.apply_chat_template( messages, tokenize=True, return_dict=True, return_tensors="pt", add_generation_prompt=True, enable_thinking=False,).to(model.device)n = inputs["input_ids"].shape[-1]out = model.generate(**inputs, max_new_tokens=64, do_sample=False)print(processor.decode(out[0][n:], skip_special_tokens=True))PY
The agent's scaffolding is a great first draft. Model names and library versions, though, I now confirm against the official model card myself — and that one habit has made this whole category of stumble disappear.
Compare adapter_config.json files before you merge
Most merge failures are knowable before you run anything. Every LoRA folder carries an adapter_config.json recording which base model it was trained on, at what rank, and on which layers. I once forgot that my translation adapter had been trained on a different base, waited through a long merge, and watched it fail at the end.
Since then, every merge starts with this preflight:
# lora_preflight.py# Purpose: read adapter_config.json from several LoRAs and decide whether they can be merged.# Usage: python lora_preflight.py ./artifacts/lora_summary ./artifacts/lora_translate ...# Exit codes: 0 = merge as-is / 1 = pick a different method / 2 = cannot merge (different base)import jsonimport sysfrom pathlib import Pathdef load(path: str) -> dict: cfg = json.loads((Path(path) / "adapter_config.json").read_text()) return { "path": path, "base": cfg.get("base_model_name_or_path"), "r": cfg.get("r"), "alpha": cfg.get("lora_alpha"), "targets": sorted(cfg.get("target_modules") or []), }def main(paths: list[str]) -> int: cfgs = [load(p) for p in paths] for c in cfgs: print(f"{c['path']}: base={c['base']} r={c['r']} alpha={c['alpha']} targets={len(c['targets'])}") if len({c["base"] for c in cfgs}) > 1: print("❌ Base models differ. Retrain on the same base or drop this adapter from the merge") return 2 code = 0 if len({c["r"] for c in cfgs}) > 1: print("⚠️ Ranks differ. linear / ties won't work; use cat or an *_svd variant") code = 1 if len({tuple(c["targets"]) for c in cfgs}) > 1: print("⚠️ Target modules differ. Layers present in only one adapter keep only that adapter's effect") code = max(code, 1) if len({c["alpha"] / c["r"] for c in cfgs if c["r"]}) > 1: print("⚠️ alpha/r ratios differ. Equal weights will not mean equal influence; compensate in weights") code = max(code, 1) if code == 0: print("✅ Base, rank, and target modules match. linear / ties are safe to try") return codeif __name__ == "__main__": sys.exit(main(sys.argv[1:]))
For methods like linear and ties, peft's add_weighted_adapter() combines the A and B matrices directly, so it expects every adapter to share the same rank. When they don't, your options are cat, which concatenates the matrices, or the svd family, which re-approximates the combined delta. If preflight returns 1, this is where you decide which combination_type the next step uses.
Implementation 1: a Weighted merge for a summary × translation model
One small success makes every later merge less confusing. Here two trained LoRAs are combined linearly at 6:4.
# merge_weighted.py# Purpose: merge the summary and translation LoRAs with weights and save a single model.# Assumes: lora_preflight.py returned 0 (same rank, same base).from pathlib import Pathfrom peft import PeftModelfrom transformers import AutoModelForMultimodalLM, AutoProcessorBASE = "google/gemma-4-E2B-it"LORA_A = "./artifacts/lora_summary" # trained beforehandLORA_B = "./artifacts/lora_translate" # trained beforehandOUT_DIR = Path("./artifacts/merged_weighted_6_4")WEIGHTS = {"summary": 0.6, "translate": 0.4}def merge_weighted(): processor = AutoProcessor.from_pretrained(BASE) base = AutoModelForMultimodalLM.from_pretrained(BASE, dtype="auto", device_map="auto") # Load several adapters by name model = PeftModel.from_pretrained(base, LORA_A, adapter_name="summary") model.load_adapter(LORA_B, adapter_name="translate") # Weighted linear combination. Switch to combination_type="cat" if ranks differ model.add_weighted_adapter( adapters=list(WEIGHTS.keys()), weights=list(WEIGHTS.values()), adapter_name="merged", combination_type="linear", ) model.set_adapter("merged") # forget this and a different adapter gets baked in merged = model.merge_and_unload() OUT_DIR.mkdir(parents=True, exist_ok=True) merged.save_pretrained(OUT_DIR, safe_serialization=True) processor.save_pretrained(OUT_DIR) print(f"✅ saved to {OUT_DIR}")if __name__ == "__main__": merge_weighted()
After merge_and_unload(), the result is saved as a plain model with no adapters attached. The serving side never needs to know adapters existed — that's the real appeal of Strategy A.
Why Weighted here? Because there are only two, and they're similar in character. Two neighboring LoRAs often reach production-grade quality with a linear blend. The moment a third joins, I switch to TIES.
I set the weights by each adapter's alpha/r and its output habits, not by how important the task feels. If preflight reports different alpha/r ratios, the adapter with the larger ratio tends to dominate, so I nudge its weight down. Comparing just three points — 0.5:0.5, 0.6:0.4, 0.7:0.3 — usually shows where the sweet spot is.
Implementation 2: a TIES merge for four LoRAs without interference
Once three or more adapters are mixed, Weighted visibly loses quality. Adapters trained on different tasks sometimes push the same parameter in opposite directions, and adding them linearly cancels both out.
TIES resolves that in three stages:
Trim: drop the smallest changes inside each adapter
Elect Sign: pick each parameter's direction by majority vote across adapters
Merge: average only the values that agree with the elected direction
In short, keep each adapter's core and add only what points the same way.
The easiest route stays inside peft. Compared with Implementation 1, only combination_type and density change.
# merge_ties_peft.py (the part that differs from merge_weighted.py)ADAPTERS = { "summary": ("./artifacts/lora_summary", 0.3), "translate": ("./artifacts/lora_translate", 0.3), "code_review": ("./artifacts/lora_code_review", 0.2), "sql": ("./artifacts/lora_sql", 0.2),}first, (path0, _) = next(iter(ADAPTERS.items()))model = PeftModel.from_pretrained(base, path0, adapter_name=first)for name, (path, _) in list(ADAPTERS.items())[1:]: model.load_adapter(path, adapter_name=name)model.add_weighted_adapter( adapters=list(ADAPTERS.keys()), weights=[w for _, w in ADAPTERS.values()], adapter_name="merged_ties", combination_type="ties", # "ties_svd" if ranks differ density=0.7, # keep the top 70% of changes majority_sign_method="total",)model.set_adapter("merged_ties")merged = model.merge_and_unload()
One caveat: peft's TIES operates on the A and B matrices separately, which is not quite the same as applying TIES to the full delta each LoRA adds to the base. When I want the more faithful version, I use mergekit, where writing base+lora_path tells it to treat the LoRA as already baked into the base.
density is the Trim threshold. At 1.0 nothing is trimmed; the lower it goes, the more only the core survives. Lower it for distant tasks, raise it for close ones.
When evaluating a TIES merge, decide up front that it doesn't have to beat each individual adapter. You're trading a little single-task performance for one model that covers several tasks. Lose sight of that and you'll tune settings forever.
Implementation 3: a dynamic-LoRA server that shares one base model
Strategy B is heavier to build, but when tasks are truly independent it gives the steadiest quality. Here's a FastAPI server that switches adapters per task. I added a /stats endpoint that counts switches, because the benchmark later depends on it.
# serve_dynamic.py# Purpose: load Gemma 4 once and switch LoRA adapters per request task.# Note: switching is exclusive, so concurrent requests are serialized with an asyncio.Lock.import asynciofrom contextlib import asynccontextmanagerimport torchfrom fastapi import FastAPI, HTTPExceptionfrom pydantic import BaseModelfrom peft import PeftModelfrom transformers import AutoModelForMultimodalLM, AutoProcessorBASE = "google/gemma-4-E2B-it"ADAPTERS = { "summary": "./artifacts/lora_summary", "translate": "./artifacts/lora_translate", "code_review": "./artifacts/lora_code_review", "sql": "./artifacts/lora_sql",}state = {"model": None, "proc": None, "current": None, "lock": asyncio.Lock(), "switches": 0, "requests": 0}@asynccontextmanagerasync def lifespan(app: FastAPI): proc = AutoProcessor.from_pretrained(BASE) base = AutoModelForMultimodalLM.from_pretrained(BASE, dtype="auto", device_map="auto") first_name, first_path = next(iter(ADAPTERS.items())) model = PeftModel.from_pretrained(base, first_path, adapter_name=first_name) for name, path in list(ADAPTERS.items())[1:]: model.load_adapter(path, adapter_name=name) model.set_adapter(first_name) model.eval() state.update(model=model, proc=proc, current=first_name) print(f"✅ loaded {len(ADAPTERS)} adapters, active={first_name}") yieldapp = FastAPI(lifespan=lifespan)class Req(BaseModel): task: str prompt: str max_new_tokens: int = 256@app.post("/generate")async def generate(req: Req): if req.task not in ADAPTERS: raise HTTPException(400, f"unknown task: {req.task}") async with state["lock"]: state["requests"] += 1 if state["current"] != req.task: state["model"].set_adapter(req.task) state["current"] = req.task state["switches"] += 1 proc = state["proc"] inputs = proc.apply_chat_template( [{"role": "user", "content": req.prompt}], tokenize=True, return_dict=True, return_tensors="pt", add_generation_prompt=True, enable_thinking=False, ).to(state["model"].device) n = inputs["input_ids"].shape[-1] try: with torch.inference_mode(): out = state["model"].generate(**inputs, max_new_tokens=req.max_new_tokens, do_sample=False) except torch.cuda.OutOfMemoryError as e: raise HTTPException(503, "GPU OOM — reduce max_new_tokens") from e return {"task": req.task, "text": proc.decode(out[0][n:], skip_special_tokens=True)}@app.get("/stats")async def stats(): r = state["requests"] or 1 return {"requests": state["requests"], "switches": state["switches"], "switch_ratio": round(state["switches"] / r, 3)}# Start: uvicorn serve_dynamic:app --port 8000
generate blocks the event loop, so in production you'd push it to a thread pool or split requests into per-task queues and process them in groups. Either way, reducing the number of switches is what decides this strategy's performance.
Ask Antigravity's agent to "add per-task queues to this server" and it will scaffold asyncio.Queue plus a time-window batcher. A scaffold is still a scaffold, though — measuring under load is the part you have to do yourself.
Benchmarks — measure switch churn against your own traffic
What I most wanted to know was how much slower things get when tasks arrive mixed together. Averages told me almost nothing. The difference shows up in P99, when switches happen back to back.
So this script sends the same number of requests two ways — grouped by task, and with the task changing every time — and prints latency percentiles alongside the switch count from /stats.
# bench_switch_churn.py# Purpose: load serve_dynamic.py while changing only the order of tasks; compare P50 / P99 / switches.# Assumes: golden_prompts.json shaped like {"summary": [...], "translate": [...], ...}import asyncioimport jsonimport randomimport timeimport httpxURL = "http://localhost:8000"TASKS = ["summary", "translate", "code_review", "sql"]PROMPTS = json.load(open("golden_prompts.json", encoding="utf-8"))def make_order(n: int, mode: str) -> list[str]: if mode == "interleaved": # worst case: the task changes every request return [TASKS[i % len(TASKS)] for i in range(n)] block = n // len(TASKS) # best case: requests arrive grouped by task return [t for t in TASKS for _ in range(block)]def pct(values: list[float], q: float) -> float: s = sorted(values) return s[min(len(s) - 1, int(q * len(s)))]async def run(mode: str, n: int = 200, concurrency: int = 8) -> None: order = make_order(n, mode) sem = asyncio.Semaphore(concurrency) lat: list[float] = [] async with httpx.AsyncClient(timeout=180) as c: before = (await c.get(f"{URL}/stats")).json() async def one(task: str) -> None: async with sem: t0 = time.perf_counter() r = await c.post(f"{URL}/generate", json={ "task": task, "prompt": random.choice(PROMPTS[task]), "max_new_tokens": 128}) r.raise_for_status() lat.append((time.perf_counter() - t0) * 1000) t0 = time.perf_counter() await asyncio.gather(*(one(t) for t in order)) wall = time.perf_counter() - t0 after = (await c.get(f"{URL}/stats")).json() switches = after["switches"] - before["switches"] print(f"{mode:12s} n={n} rps={n / wall:5.2f} " f"p50={pct(lat, .50):6.0f}ms p99={pct(lat, .99):6.0f}ms switches={switches}")async def main() -> None: random.seed(7) for mode in ("grouped", "interleaved"): await run(mode)if __name__ == "__main__": asyncio.run(main())
Reading it is simple. If the P99 gap between grouped and interleaved is small, Strategy B is fine as-is. If the gap is large and your real traffic looks more like interleaved — say, a single screen that calls several tasks in alternation — add per-task queues or move toward a Strategy A merge.
To compare against Strategy A, run the same golden_prompts.json through the merged model and line up quality with ROUGE-L or chrF++. Strategy A tends to win on speed and Strategy B on quality. Whether that gap is acceptable for your product is the whole decision.
The principle I took away: choose the strategy from the order your real requests arrive in, not from a benchmark average. The same four adapters can call for different answers under different traffic.
When Strategy C is worth it
I won't go deep here, but I'd only consider it when all three of these hold:
Task boundaries shift sentence by sentence inside a single response
A drop in single-task accuracy directly harms the people using the product
You have the time and energy to keep running a specialized inference server
S-LoRA and Punica apply multiple adapters within one batch, which makes them resilient to shifting task mixes. In exchange, GPU memory and operational effort both go up.
For solo work that's usually too heavy. I chose not to go there, and settled on Strategy A with a monthly review of the TIES settings. Not chasing the newest option is a perfectly respectable decision when what you care about is keeping things running.
Five pitfalls I actually hit
Forgetting set_adapter() before merge_and_unload(). If the adapter you built with add_weighted_adapter isn't active, whichever adapter was active gets baked in instead. If eval scores lean suspiciously toward one task, check this first.
Asking for linear or ties with mismatched ranks.peft rejects that combination. Preflight catches it early; switch to cat or an *_svd variant.
Leaving density unspecified in mergekit. Depending on your settings, a "TIES" merge can trim almost nothing and behave like a linear blend. I set density explicitly every time.
Saving in the wrong dtype. Gemma models are often trained and served in bfloat16, and dropping to float16 on save can garble outputs. Load with dtype="auto" and save in the original type.
Calling merge_and_unload() inside a dynamic-switching server. The instant you do, the adapter is baked into the base and switching stops working. I leave a comment in the server code saying never to call it there.
Observability and rollback
"Output quality quietly slipped one day" incidents are common after merges and switches. The model didn't change; the input distribution did, drifting out of the range the adapter was good at.
At minimum I watch three things:
Per-task median response length and latency. A sudden change in length — much shorter, or starting to repeat phrases — is a sign the adapter isn't doing its job.
Golden-prompt regression diffs. 30–50 evaluation prompts per task, compared against the previous release with ROUGE-L or chrF++. The benchmark's golden_prompts.json doubles as this set.
User-side signals. Regenerate-button rate, or how often a translation gets hand-edited, tend to move before the quantitative metrics do.
Rollback should be nearly instant: swap a tagged safetensors file for a merged model, or repoint a symlink to the adapter directory for dynamic switching. Just knowing I can undo a release in seconds has made the nights after an update noticeably calmer.
What I'd ask you to try this week
Run your existing adapters through lora_preflight.py first. Then take two with the same base and rank, and build exactly two Weighted merges — 0.5:0.5 and 0.7:0.3 — and compare them. That alone gives you most of the feel for merging.
TIES and dynamic switching sink in faster once that first small step is behind you. Finishing the first comparison on a weekend and starting benchmarks the following week turned out to be the right pace for me.
For the groundwork of running Gemma 4 locally, the Gemma 4 local-LLM production guide covers the surrounding infrastructure this article deliberately sets aside.
Share
Thank You for Reading
Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.