Our services

If you can think it, we can make it brainsoft.

Adding an LLM feature to an existing app without a rewrite

Written By: BrainSoft In AI & Research

Someone on the product side has seen a demo and now wants the app to summarise tickets, or draft replies, or pull structured data out of messy PDFs. The app is three years old, has real users, and nobody wants to hear the word "rewrite". Fair. You do not need one. You need a boundary.

Add the LLM as a separate service behind one narrow interface, call it from a single place in the existing code, and keep every prompt, timeout and fallback inside that boundary. The old app keeps working exactly as before; the new capability is opt-in, observable, and can be switched off with a config flag.

Find the one seam in your codebase

Every feature request sounds like it touches everything. It usually touches one thing. Ask the question backwards: if this capability disappeared tomorrow, which code path would lose it? That path is your seam. For a ticket summariser it is the function that renders the ticket detail page. For a document parser it is the upload handler. For a reply drafter it is the compose form's submit.

Write that seam down before you write any code. If the answer is "about nine places", stop and pick the highest-value one. Ship that first. The second integration is always cheaper once the shape exists.

Put the model behind a service, not inside your controllers

The temptation is to call the provider SDK straight from the route handler. It works, and it ages badly. Provider APIs change shape, keys rotate, and you will want to swap models per feature. Wrap it once.

// llm/summarise.js
export async function summariseTicket({ subject, body }) {
  const prompt = buildPrompt(subject, body);
  try {
    const raw = await client.complete(prompt, { timeoutMs: 8000 });
    return { ok: true, summary: raw.trim() };
  } catch (err) {
    log.warn({ err, feature: 'ticket_summary' }, 'llm call failed');
    return { ok: false, summary: null };
  }
}

Note the return shape. Callers never see a thrown provider error, and they never see a raw string they have to trust. They get a result object. That one decision keeps the failure mode contained.

Treat prompts like config and calls like network I/O

Two habits prevent most of the pain later.

  • Keep prompts in their own files, versioned, with the model name next to them. When output quality shifts, you want to know which prompt and which model produced it.
  • Assume every call can be slow, fail, or return something unexpected. Set an explicit timeout. Validate the response before it reaches your database or your UI.
const result = await summariseTicket(ticket);
if (!result.ok) {
  return renderTicket(ticket, { summary: null });
}
return renderTicket(ticket, { summary: result.summary });

The page renders either way. The summary is decoration, not a dependency. That is what makes the rollout boring, and boring is the goal.

Roll it out behind a flag and watch it for a week

Gate the new path on a flag keyed by user or tenant. Turn it on for staff first. Log latency, token counts and failure reasons from day one, because the first question you will get is "how much is this costing us", and you want the number already sitting in your dashboard.

If you want a second pair of eyes on the boundary before it ships, or you need the logging and hosting side handled, our services cover that. The point is that none of this required touching the parts of the app that already work.

Frequently asked questions

Do I need to move to a new framework or database?

No. The LLM feature sits behind a function call. Your existing stack, schema and deployment stay as they are. The only new pieces are a service module, a config flag and somewhere to log calls.

What if the model returns something wrong or off-topic?

Validate before you persist or display. For summaries, showing the text next to the original source lets a human catch nonsense. For structured extraction, parse into a schema and reject anything that does not fit rather than writing it to the database.

Should I self-host a model or call a hosted API?

Start hosted. It removes GPU capacity planning from a project that is really about product behaviour. Move to self-hosting when you have a concrete reason, such as data residency rules or steady volume where the per-call cost clearly outweighs running your own inference.


#AI & Research