AI productivity platform
Montlify
One workspace holding 33+ AI tools for writing, research and chat, with the accounts, subscriptions and infrastructure needed to keep them all running.
- React
- LLM APIs
- Streaming
- Authentication
Bench · AI/ML engineers
Engineers who have taken an LLM feature past the demo — through evaluation, latency, cost and the failure modes that only appear with real users.
$34 – $52 / hr · from $5,400 / month
If more than one is true, the problem is usually capacity rather than skill — and that is a different conversation.
Written as work, not as keywords. If your problem is not on this list, say so on the call — we will tell you honestly whether it is ours.
Retrieval pipelines, tool-using agents, structured output, prompt and context management. We build the surrounding system, which is where almost every LLM feature actually fails.
A graded eval set, regression gates in CI, and a number your team agrees means 'better'. Without this, every model change is a guess with a confident tone.
LoRA and full fine-tunes, distillation to smaller models, and an honest recommendation when a prompt change would have done the job for a tenth of the cost.
Detection, segmentation, OCR and document understanding, taken to a latency and cost budget you set in advance.
Quantisation, batching, caching, GPU scheduling and the serving path. Typically where the largest cost reduction in an AI product is hiding.
Training and deployment pipelines, feature and model registries, drift monitoring, and rollback that works at 2am.
We work inside what you already have. Nothing on this list is a recommendation to migrate.
Models & frameworks
Providers
Retrieval
Ops
Full-time, one engineer, includes ML review by our PhD lead.
| Role | Seniority | Rate |
|---|---|---|
| AI/ML Engineer | Mid-senior · 3–5 yrs | $34 – $42 / hr |
| Senior AI/ML Engineer | Senior · 5–8 yrs | $42 – $52 / hr |
| ML Research Engineer | PhD-level | $58 – $75 / hr |
| LLM Platform Engineer | Senior · 5+ yrs | $45 – $55 / hr |
Rates are per hour for full-time engagement and include our management, review and replacement guarantee. Part-time and pod pricing differ — see services.
Working to a fixed budget, or need part-time, long-term or several engineers? Those price below the card.
Contact for pricingEvery role, every rate, our standard MSA and NDA, and the two-week trial terms. No call required to read it.
No two-week onboarding. The first week produces something you can read.
Answer accuracy on a client RAG system, without changing model
Inference cost reduction on a production vision workload
Median time from contract signed to first merged PR
The same five steps regardless of discipline.
You describe the gap. We tell you which discipline it actually is — that answer changes about a third of the time — and whether we have the person.
Real engineers with availability, not a database dump. Each one comes with a written note on why they fit and where they would struggle.
Your process, your bar, your rejection. We do not present anyone we would not hire ourselves, and the engineer who interviews is the engineer who ships.
They join your standups and open real pull requests. Stop inside the trial for any reason and the engagement ends there.
Monthly rolling from there. A senior lead reviews their work independently of you, and we tell you before you have to ask.
AI productivity platform
One workspace holding 33+ AI tools for writing, research and chat, with the accounts, subscriptions and infrastructure needed to keep them all running.
Yes. We work inside whatever you have — Anthropic, OpenAI, Bedrock, Vertex, or self-hosted weights on your own GPUs. We do not resell inference and have no incentive to move you.
Regularly. Engineers work in your cloud accounts, on your machines or ours under your policy, with signed NDAs and per-engineer access scoping. Several of our engagements run entirely inside a client VPC.
You'll get that answer, in week one, in writing. It has happened. It costs us a month of billing and buys a client who comes back.
Most teams start with one engineer and add a second from a neighbouring bench within two quarters.
Service and API engineers who have carried a pager, and who write the boring, well-tested code that keeps you off one.
Pipeline and warehouse engineers who care whether the number in the dashboard is correct, and can prove it.
Infrastructure engineers who reduce the number of things that can page you, and who treat your cloud bill as an engineering artefact.
Thirty minutes, an engineer on the call, and two or three real profiles within five working days.