A person stays in charge
Local Orchestration and Custom LLMs for Agent Workforces
If your agent army only exists inside a single vendor chat UI, you do not have an operating model — you have a subscription.
MeltingFace researches local orchestration: agent runtimes, adapters, queues, and model servers you can inspect on infrastructure you control (or at least policy you can name). Custom LLMs sit next to that stack — not as a brand stunt, but as a way to fit house tools, house style, and house context floors.
What “local orchestration” means
Local does not mean “never use the cloud.” It means:
- Control plane on the host — services, profiles, hybrid wrappers, issue orchestration
- Explicit model endpoints — local inference servers with known context length, layer offload, and restart policy
- Adapters with audit — agent processes that record spawn, timeout, and disposition
- Fail closed — when the local model dies, the company does not silently open a paid API with agent spend authority
For agent-first companies, that control plane is what lets humans stay in charge while agents run for hours.
Why custom / adapted LLMs matter
Off-the-shelf models are excellent generalists. Company work is not general:
- Long context floors for real tickets and specs
- Tool schemas and house CLI conventions
- Brand voice that must not invent legal identity
- Coding and review styles that match your repo
Adaptation paths (fine-tuning, continued pretrain, preference optimization, system-prompt + tool training packs) are research tools for fitting models to that reality. Quantization and partial GPU offload are operational tools for running them next to a desktop and a display GPU without pretending the machine is a full data center.
We talk about these choices openly because agent productivity collapses when the model thrashing the GPU is misconfigured for the job.
Orchestration layers we care about
- Company graph — issues, parents, assignees, terminal states
- Agent runtime — profiles, timeouts, max iterations, sole-leaf gates
- Model server — context, batch, offload, recovery after device loss
- Policy — dry-run, no real money, brand kit read-only
- Human gate — Board unlock, Janus review, content status
Skip a layer and the army either idles or burns.
What we are not claiming
- That every company should train a foundation model from scratch
- That local always beats remote on quality
- That any specific tok/s number is a market ranking without a published protocol
We claim that orchestration + model fit + gates is the productive unit of research — not the model name alone.
Read next
*Engineering research notes. Configurations and commercial offerings are Board-gated; no live multi-tenant claims in v0.1.*
MeltingFace