Skip to content
THE GUILD
0%
SERVICES PRODUCTS CAREERS ABOUT BLOG FAQ CONTACT

Private LLM PoC · 30 days, fixed price

Prove a private LLM works on your data, in 30 days — before you commit to hardware.

Accuracy and latency measured against your own documents and workload, not a generic benchmark. You get a written recommendation on deployment form, model size, and full-build cost — before spending on a GPU server.

Start a Private LLM PoC Free initial call first · NDA before any data changes hands · Reply within 1 business day

Common situations

"We want a private LLM. We do not want to spend on a GPU server to find out if it actually works."

A Private LLM PoC is for the moment before hardware or a hosting contract gets committed — when the open question is still "does this meet our accuracy bar," and you want an answer on your own data, not a vendor's demo.

Situation 01

A full 8-GPU build runs into six figures, in USD. That is a lot to commit to before anyone has measured accuracy on your own documents.

See the equipment-market reality on the infrastructure service page — this PoC exists so you do not have to spend that first.
Situation 02

IT already tried self-hosting an open model. It either fell short of the accuracy people are used to, or needed more VRAM than expected.

Model size and quantization are exactly what the PoC tests against your real workload.
Situation 03

Self-hosted, managed, or shared — three ways to run this, and no one internally can say with confidence which is cheapest for your usage pattern.

None of the three wins at every scale. The PoC report says which one, for you.

Before → After

What changes in 30 days.

The deliverable is a decision you can act on, not a slide deck about what a private LLM could theoretically do.

Day 0 · Before
  • No hardware, and no accuracy numbers measured against your own documents.
  • Three deployment forms on paper — self-hosted, managed, shared — with no basis to choose one.
  • A GPU market moving fast enough that last quarter's quote may already be stale.
  • No estimate for what a full build would actually cost for your specific use case.
Day 30 · After
  • A written report: feasible, conditionally feasible, or not feasible — for your workload specifically.
  • Accuracy and latency measured against your own documents and questions.
  • A recommended model tier and deployment form, with the reasoning behind each.
  • A cost estimate for the full build, and the PoC fee credited toward it if you continue.

Why this PoC, before a hardware quote

The point is a decision you can defend, not a demo.

A private LLM PoC is built to answer one question honestly: does this meet your accuracy and latency bar, on your own material, at a cost that makes sense for your usage.

01 · No hardware required

Tested on infrastructure we already operate

The PoC runs on LLL's own in-house environment. You are not asked to provision a GPU server to find out whether the answer is yes.

02 · Your data, not a benchmark

Accuracy measured on your own documents

The evaluation set is built from your material and your questions — not a generic public benchmark that may not reflect your domain or your language use.

03 · A recommendation, not a pitch

Deployment form and model size, with reasoning

The report names a specific model tier and deployment form and explains why — including saying so if a private LLM is not the right fit for what you tested.

How it works · 4 weekly phases

Thirty days from kickoff to a written recommendation.

NDA and DPA are signed before any of your documents or data are used in evaluation.

Week 1

Kickoff & workload definition

Define the use case, the documents and questions the model will be tested against, and what "feasible" means for your accuracy and latency requirements — in writing, before testing starts.

NDA/DPA signed first

Week 2

Shortlist model & deployment form

Based on the defined workload, shortlist a model tier (starting from the 70B-class default) and provision it on our own operating environment for testing.

No hardware purchase at this stage

Week 3

Build & test

Connect the shortlisted model to your material (RAG where relevant) and run accuracy tests against your own evaluation set, plus latency and throughput under a representative load.

Results measured, not estimated

Week 4

Report & recommendation

Written validation report: feasible, conditionally feasible, or not feasible. Recommended model tier and deployment form, plus a cost estimate for the full build.

Full-build fee credited if you continue

What the PoC actually removes

The commitment you are not making yet.

These are the three things a private LLM PoC is designed to remove before you sign anything larger.

0Hardware purchased for the PoC

Testing runs on infrastructure LLL already operates. You commit to hardware only after the report tells you which form and size actually fit.

FP8Quantization floor, tested as-is

The same floor we use in every recommended build — our testing shows a measurable accuracy drop on Japanese-language business tasks below it, so we do not test or ship lower by default.

3Deployment forms compared

Self-hosted, managed, and shared — each has a scale where it is cheapest. The report says which one for your workload, not by default.

What you can rely on

Five rules that govern every Private LLM PoC.

Operational rules, not stretch goals — the same on every engagement.

  • NDA before data

    NDA and DPA are signed before any of your documents, data, or questions are used for evaluation — including during the free initial call.

  • Tested on your own material

    Accuracy is checked against your documents and questions, not a generic public benchmark that may not reflect your domain or language use.

  • FP8 quantization floor

    We do not test or recommend below FP8 for business-facing answers. If a use case genuinely does not need it, the report says so — it is not an upsell.

  • Full credit toward the build

    If you continue to a full build, the $6,000 (or $3,000) PoC fee is deducted from that invoice in full. No voucher, no expiry.

  • Honest no-go

    If the workload does not fit a private LLM — accuracy falls short, the economics do not work, a public API is genuinely the better fit — the report says so instead of recommending a build anyway.

FAQ

Five questions before you start.

Specific case not here? Add it to the message field on the form below. We reply within one business day.

What hardware does the PoC actually run on?
The PoC runs on infrastructure LLL already builds and operates in-house. You are not asked to buy a server to find out whether a private LLM fits your use case — the report at the end tells you what to buy, rent, or have us host, if you decide to continue.
Which model size will you test?
We start from the recommended default — a 70B-class dense model (about 70GB in FP8), sized to approximate a company-wide ChatGPT-equivalent on one machine. If your workload needs less (the 32B floor) or more (122B-A10B or 400B-class MoE for heavy concurrent usage), the report says so with reasoning, not a default upsell.
Why FP8 and not a smaller quantization?
Our standard configurations hold FP8 as the floor. Quantizing further (for example, 4-bit) measurably degrades accuracy on Japanese-language business tasks in our testing, so we do not ship it as a default recommendation even though it would lower the hardware bill.
Does the fee really come off the full build?
Yes. The full $6,000 (or $3,000 for teams of 50 or fewer) is deducted from the cost of the full build if you continue. No voucher, no expiry — it comes off the next invoice.
Does "private" mean nothing ever leaves our network?
It means inference — the prompt and the generated answer — is processed on servers inside your own infrastructure rather than sent to a public, general-purpose AI service. It is not a claim about your network perimeter, VPNs, or access controls, which remain a separate design decision from where the model itself runs.

Start a Private LLM PoC

Tell us the workload. Get a free initial assessment.

We reply within 1 business day with a free initial read on feasibility, before you commit to the $6,000 (or $3,000).

  • NDA signed before any technical discussion — including the free assessment.
  • No hardware purchase required to run the PoC itself.
  • If a private LLM is not the right fit, we will say so in the report.
  • $6K ($3K ≤50 staff)
  • 30 days
  • 0 hardware required
  • Credit toward full build

Request a free initial assessment

Required fields marked *

By submitting, you agree to our privacy notice.