LumaBrowser 1.7 is a series of nine point releases. Half of it is the chat agent learning to call tools natively, keep long conversations within budget, and never edit a file it has not read. The other half is the model layer: two new runtimes, a smarter multi-GPU split, RAM pinning, and a benchmark that tells you whether a model actually behaves, not just whether it loads.
Local models are good at tool use only when the plumbing does not get in their way. Most of 1.7's chat work is plumbing.
Models that speak a native tool-call format (the Qwen family first) now use it instead of a text fence, with the fence kept for models that do not. Mixed and malformed call shapes are normalised, illegal JSON escapes recovered, and a call is no longer lost to a single backslash or a nested code block.
Long conversations are summarised against a cache-warm prefix instead of being cut off, so the model keeps the thread without reprocessing everything. One owner for the context budget, measured per-request fixed cost, and history budgeted against the usable window rather than the whole one.
The agent loop can issue several tool calls in one turn. A second state-changing call in the same batch is deferred, and the default pool size is set from benchmark evidence rather than optimism.
Before a tool changes anything on disk or on the web, the chat asks once. Tool cards now show the call separately from the model's prose, and required arguments are enforced at dispatch on both transports.
The most recent assistant reply can be edited inline, with save and cancel, so a nearly-right answer does not need a whole regeneration.
Thinking models get a per-conversation effort control, and the agent loop gets a tool loop's reasoning budget instead of a chat reply's. A provider-confirmed context overflow is recovered once instead of ending the run.
Code Mode refuses to edit a file the agent has not read in the current context, and a stale-write guard rejects an edit against a version that has changed underneath it. Files are identified by content, not by their stat.
Workspace search runs on the packaged ripgrep with a directory walk as fallback, and discloses which directories it skipped.
Describe a game and get a playable one, with sprites and backdrops generated by the image server (pixel-art native sizes, nearest-neighbour upscaling, a batch asset tool that plays nicely with hotswap). A headless runner reports runtime errors and phantom API calls before you see them.
Games can call the model at runtime through a small bridge, keep persistent data stores, and relay a multiplayer room. Design guidelines for core loops, difficulty pacing, and genre recipes ship with the mode.
What the agent read on the web is previewable from the chat, so you can see the page it used rather than trust the summary.
The Anthropic provider handles max_tokens and reasoning for adaptive-thinking models, so an Opus-class model no longer loops silently under a token cap.
Every one of these is a knob on the advanced page. Easy Setup uses them for you.
ik_llama (1.7.7) as a data-only add-on on the Thireus prebuilt feed, ISA-matched to your CPU, with its native quant patterns. NInfer (1.7.6) as an add-on runtime with add-on model support, measured at roughly twice llama.cpp's Q6 speculative decode on our rig.
A CUDA 12 build of llama.cpp with our patch stack on top of upstream: faster mmap load (about 3x on the weight phase) and the multi-GPU fixes below. Selectable next to upstream in the runtime list.
A plain two-GPU split now fills the fastest card first instead of splitting evenly. On a 5090 + 3090 at 128k context that is the difference between 45.8 and 64.9 tokens per second, because the KV no longer spills into WDDM shared memory.
Chat and image models can be locked into RAM so a large mmap'd model does not page out between requests. The planner chooses no-mmap versus mmap per launch shape, preflight is cached, and eviction overlaps the launch (5.6 s to 0.1 s on a relaunch). Fixed for packaged builds in 1.7.8 and 1.7.9.
Model-free n-gram speculation with a draft cap (measured +26 percent on edit-heavy turns for Qwen 3.8 Flash-Next), plus detection of MTP-grafted quants. The launch planner also auto-tunes for memory fit and predicts parallel slots.
A behavioural integration test for the model you launched: tool calling, editing, and long-context tasks, scored with the context and GPUs the run actually used. Sits beside the fit test in the Setup tab and runs from the command line for A/B diffs that include wall time.
Multi-card VRAM claims are weighted and carry a resident flag, the text encoder for image models is placed on CPU or GPU by a decision rather than a hard-coded flag, and the health poll dropped from 1000 ms to 200 ms so servers report ready sooner.
LoRA files are inspected for their base model before attaching. Catalog additions include Chroma1-HD, FLUX.2 klein 4B, Sprite Shaper XL, and pixel-art and seamless-texture LoRAs for game assets.
Model references moved to Qwen 3.8 (1.7.3) with per-family defaults: the text-fence tool protocol by default, edit_artifact kept off the native tools array, and size-weighted tool-history eviction.
These change behaviour you may be relying on. Read this section before upgrading a machine that other people or scripts depend on.
It shipped behind a flag, default off, and was then turned on for models that support it. If you parse the model's raw output for a text fence, expect native tool_calls on those models instead. Inactive tool groups are stubbed natively.
The one-shot approval prompt applies to chat, sub-agents, and scheduled runs. Unattended runs need the tool pre-approved in the chat's tool settings or they will wait.
An edit against a file the agent has not read, or against a stale version, is rejected with a reason. Scripts that drove edits through the agent without a preceding read will need one.
Parallel tool calls exist but the default pool size is one, set from gambit evidence. Raise it per model family if your model handles it.
The bytecode packer was not unpacking the native module and worker scripts those features need. Fixed in 1.7.9; if you are on an earlier 1.7 build and pinning silently did nothing, update.
Family sampler defaults changed again (penalty handling), so generated text will not match previous runs byte for byte.
These ship in 1.7 but have not finished hardening. They are listed so you know they exist and what their limits are.
The automatic image runtime install does not work on Linux in 1.7.9: the installer cannot unpack the archive it downloads, and on NVIDIA it picks a CUDA build that has no Linux version. Chat is unaffected. The image guide has a manual workaround; the fix is in progress.
Builds and plays real games with generated assets, verified end to end on Qwen 3.8. A mid or large chat model matters a lot here; small models produce small games.
Both install and run. Neither has the years of edge-case coverage upstream llama.cpp has; bench them with the gambit before making one your default.
Speech in and out using local models, Windows and Linux only; there is no whisper binary release for macOS yet.
Present but not exercised widely. Windows requires WSL2, macOS is not supported, and the model wants about 34 GB of VRAM on a single card.
Off by default. The transport carries no authentication and no encryption; use it only between machines you own on a network you control.
1.7.9 inline edit of the last reply; bytecode unpack fix for RAM pinning and TTS; version bump script.
1.7.8 LumaByte llama.cpp CUDA 12 build in the catalog; memory-fit auto-tuning; 200 ms health poll; dynamic RAM pinning; LoRA inspector; n-gram speculation; sprite and backdrop model selection; CLIP placement decision; VRAM coordinator weighted claims; AI probe for games; pixel-art native sizes; fill-order layer split.
1.7.7 ik_llama runtime.
1.7.6 NInfer runtime with add-on models.
1.7.5 persistent data stores for AI-driven games.
1.7.4 adaptive thinking on Anthropic; game design guidelines; sandboxed batch asset generation; headless game runner; Chroma1-HD, FLUX.2 klein, Sprite Shaper XL.
1.7.3 Qwen 3.8 as the reference family; multiplayer room relay; game asset generation.
1.7.1 chat compaction; parallel tool calls; read-before-edit and stale-write guard; native tool calling; one-shot approval; packaged ripgrep search; context budget owner.
1.7.0 compatibility gambit; reasoning effort control; web usage preview.
LumaBrowser is free to download and runs local models with no account and no API key. Read the API reference if you want to drive it from your own code.