Release notes · Version 1.8 · current build 1.8.1

The model got better at using tools, and the hardware got easier to use.

LumaBrowser 1.7 is a series of nine point releases. Half of it is the chat agent learning to call tools natively, keep long conversations within budget, and never edit a file it has not read. The other half is the model layer: two new runtimes, a smarter multi-GPU split, RAM pinning, and a benchmark that tells you whether a model actually behaves, not just whether it loads.

Download LumaBrowser - free Changes to know about
Windows · macOS · Linux · version list

A chat agent that holds up over a long session

Local models are good at tool use only when the plumbing does not get in their way. Most of 1.7's chat work is plumbing.

Native tool calling

Models that speak a native tool-call format (the Qwen family first) now use it instead of a text fence, with the fence kept for models that do not. Mixed and malformed call shapes are normalised, illegal JSON escapes recovered, and a call is no longer lost to a single backslash or a nested code block.

Chat compaction

Long conversations are summarised against a cache-warm prefix instead of being cut off, so the model keeps the thread without reprocessing everything. One owner for the context budget, measured per-request fixed cost, and history budgeted against the usable window rather than the whole one.

Parallel tool calls, with a brake

The agent loop can issue several tool calls in one turn. A second state-changing call in the same batch is deferred, and the default pool size is set from benchmark evidence rather than optimism.

One-shot approval

Before a tool changes anything on disk or on the web, the chat asks once. Tool cards now show the call separately from the model's prose, and required arguments are enforced at dispatch on both transports.

Edit the last reply

The most recent assistant reply can be edited inline, with save and cancel, so a nearly-right answer does not need a whole regeneration.

Reasoning effort per conversation

Thinking models get a per-conversation effort control, and the agent loop gets a tool loop's reasoning budget instead of a chat reply's. A provider-confirmed context overflow is recovered once instead of ending the run.

Building things: safer edits, playable games

Read before edit

Code Mode refuses to edit a file the agent has not read in the current context, and a stale-write guard rejects an edit against a version that has changed underneath it. Files are identified by content, not by their stat.

Packaged search

Workspace search runs on the packaged ripgrep with a directory walk as fallback, and discloses which directories it skipped.

Game mode with generated assets

Describe a game and get a playable one, with sprites and backdrops generated by the image server (pixel-art native sizes, nearest-neighbour upscaling, a batch asset tool that plays nicely with hotswap). A headless runner reports runtime errors and phantom API calls before you see them.

AI-driven games and persistence

Games can call the model at runtime through a small bridge, keep persistent data stores, and relay a multiplayer room. Design guidelines for core loops, difficulty pacing, and genre recipes ship with the mode.

Web usage preview

What the agent read on the web is previewable from the chat, so you can see the page it used rather than trust the summary.

Adaptive thinking on Anthropic

The Anthropic provider handles max_tokens and reasoning for adaptive-thinking models, so an Opus-class model no longer loops silently under a token cap.

Models and hardware

Every one of these is a knob on the advanced page. Easy Setup uses them for you.

Two new runtimes

ik_llama (1.7.7) as a data-only add-on on the Thireus prebuilt feed, ISA-matched to your CPU, with its native quant patterns. NInfer (1.7.6) as an add-on runtime with add-on model support, measured at roughly twice llama.cpp's Q6 speculative decode on our rig.

LumaByte llama.cpp build

A CUDA 12 build of llama.cpp with our patch stack on top of upstream: faster mmap load (about 3x on the weight phase) and the multi-GPU fixes below. Selectable next to upstream in the runtime list.

Fill-order layer split

A plain two-GPU split now fills the fastest card first instead of splitting evenly. On a 5090 + 3090 at 128k context that is the difference between 45.8 and 64.9 tokens per second, because the KV no longer spills into WDDM shared memory.

Dynamic RAM pinning

Chat and image models can be locked into RAM so a large mmap'd model does not page out between requests. The planner chooses no-mmap versus mmap per launch shape, preflight is cached, and eviction overlaps the launch (5.6 s to 0.1 s on a relaunch). Fixed for packaged builds in 1.7.8 and 1.7.9.

N-gram speculation and MTP

Model-free n-gram speculation with a draft cap (measured +26 percent on edit-heavy turns for Qwen 3.8 Flash-Next), plus detection of MTP-grafted quants. The launch planner also auto-tunes for memory fit and predicts parallel slots.

The compatibility gambit

A behavioural integration test for the model you launched: tool calling, editing, and long-context tasks, scored with the context and GPUs the run actually used. Sits beside the fit test in the Setup tab and runs from the command line for A/B diffs that include wall time.

VRAM coordinator and CLIP placement

Multi-card VRAM claims are weighted and carry a resident flag, the text encoder for image models is placed on CPU or GPU by a decision rather than a hard-coded flag, and the health poll dropped from 1000 ms to 200 ms so servers report ready sooner.

LoRA inspector and new image models

LoRA files are inspected for their base model before attaching. Catalog additions include Chroma1-HD, FLUX.2 klein 4B, Sprite Shaper XL, and pixel-art and seamless-texture LoRAs for game assets.

Qwen 3.8 as the reference family

Model references moved to Qwen 3.8 (1.7.3) with per-family defaults: the text-fence tool protocol by default, edit_artifact kept off the native tools array, and size-weighted tool-history eviction.

Changes worth knowing about

These change behaviour you may be relying on. Read this section before upgrading a machine that other people or scripts depend on.

Native tool calling is on for the Qwen family

It shipped behind a flag, default off, and was then turned on for models that support it. If you parse the model's raw output for a text fence, expect native tool_calls on those models instead. Inactive tool groups are stubbed natively.

A tool that changes state asks first

The one-shot approval prompt applies to chat, sub-agents, and scheduled runs. Unattended runs need the tool pre-approved in the chat's tool settings or they will wait.

Code Mode will refuse a blind edit

An edit against a file the agent has not read, or against a stale version, is rejected with a reason. Scripts that drove edits through the agent without a preceding read will need one.

Default tool-call pool is 1

Parallel tool calls exist but the default pool size is one, set from gambit evidence. Raise it per model family if your model handles it.

Packaged builds: RAM pinning and TTS were dead in 1.7.7 and 1.7.8

The bytecode packer was not unpacking the native module and worker scripts those features need. Fixed in 1.7.9; if you are on an earlier 1.7 build and pinning silently did nothing, update.

Local model output will differ on identical inputs

Family sampler defaults changed again (penalty handling), so generated text will not match previous runs byte for byte.

Early access in this release

These ship in 1.7 but have not finished hardening. They are listed so you know they exist and what their limits are.

Experimental

Linux image generation

The automatic image runtime install does not work on Linux in 1.7.9: the installer cannot unpack the archive it downloads, and on NVIDIA it picks a CUDA build that has no Linux version. Chat is unaffected. The image guide has a manual workaround; the fix is in progress.

Game mode

Builds and plays real games with generated assets, verified end to end on Qwen 3.8. A mid or large chat model matters a lot here; small models produce small games.

NInfer and ik_llama runtimes

Both install and run. Neither has the years of edge-case coverage upstream llama.cpp has; bench them with the gambit before making one your default.

Voice conversation mode

Speech in and out using local models, Windows and Linux only; there is no whisper binary release for macOS yet.

Music generation

Present but not exercised widely. Windows requires WSL2, macOS is not supported, and the model wants about 34 GB of VRAM on a single card.

Borrowing GPUs from another machine

Off by default. The transport carries no authentication and no encryption; use it only between machines you own on a network you control.

The 1.7 series

1.7.9 inline edit of the last reply; bytecode unpack fix for RAM pinning and TTS; version bump script.
1.7.8 LumaByte llama.cpp CUDA 12 build in the catalog; memory-fit auto-tuning; 200 ms health poll; dynamic RAM pinning; LoRA inspector; n-gram speculation; sprite and backdrop model selection; CLIP placement decision; VRAM coordinator weighted claims; AI probe for games; pixel-art native sizes; fill-order layer split.
1.7.7 ik_llama runtime.
1.7.6 NInfer runtime with add-on models.
1.7.5 persistent data stores for AI-driven games.
1.7.4 adaptive thinking on Anthropic; game design guidelines; sandboxed batch asset generation; headless game runner; Chroma1-HD, FLUX.2 klein, Sprite Shaper XL.
1.7.3 Qwen 3.8 as the reference family; multiplayer room relay; game asset generation.
1.7.1 chat compaction; parallel tool calls; read-before-edit and stale-write guard; native tool calling; one-shot approval; packaged ripgrep search; context budget owner.
1.7.0 compatibility gambit; reasoning effort control; web usage preview.

Run all of it on your own machine

LumaBrowser is free to download and runs local models with no account and no API key. Read the API reference if you want to drive it from your own code.