LM Studio Bionic Splash Engine: Run Qwen3.8 Faster on Apple Silicon in 2026
LM Studio Bionic now has a Splash inference engine for Apple Silicon, aimed specifically at fast local Qwen3.8 and Qwen3.6 inference. This is a much narrower feature than a general “faster LM Studio” update: it targets particular Macs and particular models.
According to LM Studio's current documentation, Splash is an open-source inference engine from Inco AI optimized for Qwen3.6-35B-A3B and Qwen3.8-27B. This guide explains exactly who can use it and how to set it up.
What Is Splash Engine?
Splash is a specialized local inference engine designed for Apple Silicon. LM Studio's September 2026 documentation describes model-specific GPU kernels and memory planning, plus dedicated draft models for speculative decoding.
Inco AI's reported testing on a 48 GB M5 Pro showed roughly 74 tokens/second on short Qwen3.8-27B prompts and 54 tokens/second at 32K context. Those are vendor-reported results, not a guarantee for every Mac.
Hardware Requirements
| Requirement | Current requirement |
|---|---|
| Mac | M3 or newer |
| macOS | 26.4 or later |
| Unified memory | At least 36 GB |
| Recommended memory | 48 GB or more |
| Bionic | 1.1.5 or newer |
This means an M1 or M2 Mac does not meet the documented Splash requirement, even though those Macs can run many other local models through LM Studio.
How to Enable Splash in Bionic
- Install LM Studio Bionic 1.1.5 or newer.
- Open Settings → Runtime.
- Find Experimental backends.
- Download Splash (Metal).
- Open Settings → Explore.
- Search for the supported model.
- Download a Splash-compatible Qwen model.
- Start a new Bionic session and select the model.
LM Studio currently documents these model repositories: incoai/Qwen3.8-27B-Splash and incoai/Qwen3.6-35B-A3B-Splash.
Why Splash Is Different
A normal local inference engine tries to support many model architectures and hardware configurations. Splash is narrower: its advantage comes from optimizing for supported models and Apple Silicon rather than trying to be a universal backend.
That trade-off matters. If you use Qwen3.8 on a supported Mac, a specialized backend may provide a significant speed improvement. If you use an unsupported model, Splash is not the reason to change your whole LM Studio setup.
Qwen3.8-27B vs Qwen3.6-35B-A3B
Both are listed by LM Studio as Splash-supported models. Qwen3.8-27B is the more obvious starting point if your goal is the current Qwen3.8 workflow. Qwen3.6-35B-A3B is another option when you specifically want that architecture/model family.
Test the same prompt, context length and workload on both models before deciding which feels better on your machine.
How to Verify That Splash Is Actually Being Used
Do not judge the backend only by the first generated response. Check the selected runtime in Bionic, confirm the model is a Splash-compatible build, and then measure a repeatable prompt.
A useful benchmark is to run the same prompt three times and compare warm-generation speed rather than a single first-token result. Record context length, model, backend and whether another workload is using the GPU.
Common Splash Problems
Splash does not appear
Check the Bionic version, Mac generation and macOS version first. The current requirements are restrictive by design.
The model is not available
Use one of the model repositories documented by LM Studio. A normal Qwen3.8 model download is not automatically a Splash build.
Performance is not what the blog claims
Vendor benchmarks use specific hardware, prompt sizes and context lengths. Your result can differ because of memory, thermal limits, background applications, model version and workload.
My Mac has enough RAM but is an older chip
Unified-memory capacity alone is not enough. Splash currently requires M3-or-newer Apple Silicon, so an older Mac should use the regular supported Bionic runtimes instead.
How This Fits With Bionic 1.1.6
Bionic 1.1.6 added Canvas, editable Markdown/source files and llama.cpp 2.43.0 extension packs. Splash was introduced in the preceding 1.1.5 release, so the features serve different purposes: Splash affects inference on supported Macs, while 1.1.6 expands the agent workspace.
For a complete workflow, you can combine Bionic's current agent features with a supported local model and use Canvas when the task involves editable code or documents.
Sources
Check the official LM Studio Splash Engine announcement and the Bionic changelog for current requirements and release changes.
Bottom Line
Splash is worth investigating if you have an M3-or-newer Mac with enough unified memory and your workload is Qwen3.8 or Qwen3.6. It is not a universal accelerator for every LM Studio model, but for its supported models it is a targeted way to push Apple Silicon local inference harder.
Performance Testing and Troubleshooting
Performance Testing and Troubleshooting
The most useful Splash test is not a benchmark number copied from another Mac. It is a repeatable workload on your own machine.
- Close unrelated heavy applications.
- Select the same Qwen model and context length for every run.
- Run one short prompt for warm decode speed.
- Run one longer prompt to test sustained generation.
- Repeat each test and record the approximate tokens/second.
- Watch memory and thermal behavior during the run.
Keep your notes with the Mac model, unified memory, Bionic version, macOS version, model build and backend. That makes future Bionic updates much easier to compare.
Splash vs Regular Bionic Runtime
| Question | Splash | Regular runtime |
|---|---|---|
| Supported hardware | M3+ with documented memory requirements | Broader hardware support |
| Model focus | Specific Qwen builds | Much broader model coverage |
| Goal | Specialized Apple Silicon speed | General local inference |
| Best reason to use | Supported Qwen workload | Mixed model collection |
So there is no need to treat Splash as a replacement for every Bionic backend. Think of it as a specialized option that can coexist with your general-purpose runtime.
Does Splash Replace Ollama or LM Studio?
No. Splash is an inference backend used within the Bionic workflow. Ollama remains useful when you want a model-serving API, CLI-based model management or integrations with coding agents. Regular LM Studio/Bionic runtimes remain useful when you need broader model compatibility.
Security and Privacy Notes
Local inference can keep prompts and files on your machine, but model downloads, external tools, browser access and remote integrations can still introduce network activity. If you are using Bionic for private source code, review the application's network and tool permissions instead of assuming that every part of the workflow is offline.
Who Should Try Splash?
- Mac users with M3-or-newer hardware.
- Systems with at least the documented 36 GB unified-memory requirement.
- Users primarily interested in Qwen3.8-27B or Qwen3.6-35B-A3B.
- Developers who value local coding-agent speed.
If your Mac does not meet those requirements, regular LM Studio/Bionic inference is the more relevant path.