New AI models arrive faster than most teams can try them. The names change, the benchmark charts get longer, and there is still a client email to write before lunch.

So what should you actually use?

We think the useful starting point is the task. Cleaning up meeting notes and preparing a recommendation from conflicting evidence need different amounts of work. The latest smaller models make the first job increasingly affordable. Stronger models are worth considering when the second job needs more reasoning, context and revision.

The quick version

  • Start with a lighter model for clearly defined work: summaries, routine drafts, extracting facts and organising information.
  • Try a stronger model when the task involves conflicting evidence, several connected steps or a deliverable that takes substantial editing.
  • Match the model to the material. A screenshot, a long PDF and an audio recording place different demands on the tools around it.
  • Compare models on something your team actually does. Include the time spent checking and correcting the answer.

Where the five models fit

These are starting points based on the providers' published capabilities, rather than results from our own benchmark. The links below lead to their official guidance.

ModelA useful place to start
Gemini 3.8 FlashLong documents and mixed material: text, images, audio and PDFs.
Claude Haiku 5.5Quick summaries, classification, lookups and repetitive support work.
Claude Opus 5.5Research synthesis, difficult analysis, substantial writing and complex coding.
GPT-6.1 SolProjects with several steps, technical work and connected professional deliverables.
GPT-6 LunaFocused edits, structured extraction and repeatable tasks with clear instructions.

There is overlap here. A smaller model can handle a difficult request when the information and instructions are clear. A larger one can still miss a simple detail. Use the table to decide what to try first.

Meeting notes, emails and everyday admin

Suppose you have a page of notes from a client call. You need a short summary, a list of actions and a follow-up email. Most of the information is already there; the model needs to organise it without inventing anything.

Haiku 5.5 and GPT-6 Luna are sensible first options. Anthropic positions Haiku for quick, repetitive workloads, including summaries and classification. OpenAI recommends Luna for focused edits and simple extraction.

The prompt matters. Ask for an owner and a deadline only where the notes name them. Ask the model to mark missing details as unknown. Give it one example of the email tone you want.

If the call contains a disagreement about scope, try Opus 5.5 or GPT-6.1 Sol to work through the competing interpretations. Separate what each person said from what the model recommends doing next.

Research and comparing documents

Imagine choosing between three suppliers. You have proposals, product screenshots and a PDF of your requirements. The first job is to extract the same facts from each source: price, delivery dates, exclusions and unanswered questions.

Gemini 3.8 Flash is worth trying here. Google documents support for text, images, audio, video and PDFs, alongside a large context window. The application you use determines which inputs and tools are exposed, so check that before building a workflow around a particular format.

Ask for a comparison with a source reference for each claim. A long context window gives the model room to read; it does not guarantee that it notices every exception.

Then comes the judgement: which proposal fits your constraints, and what might change that recommendation? Opus 5.5 and GPT-6.1 Sol are useful candidates for this stage. Give them the original evidence as well as the extracted facts, so the analysis can catch an omission in the first pass.

Client proposals and substantial writing

A routine email needs a clear message. A proposal has more moving parts: the client's problem, your approach, the scope, the evidence and the next decision.

Try Opus 5.5 or GPT-6.1 Sol when those parts need to fit together. Anthropic highlights Opus's knowledge work and communication improvements. OpenAI describes Sol as a model for complex professional work, including deliverables built across several steps.

Start with the argument and outline. Supply the facts, the audience and examples of your own writing. Once the structure works, use a lighter model for local edits, shorter versions or extracting a checklist.

We would judge the result by how much useful editing remains. A fluent page full of vague claims can take longer to fix than a rough draft with the right substance.

Coding and technical problem-solving

For a small, well-defined change, Luna or Haiku may be enough. Give the model the relevant code, the expected behaviour and a way to check the result.

Sol and Opus are stronger starting points for problems spread across several files, unfamiliar systems or ambiguous requirements. Both providers emphasise complex coding work. Gemini 3.8 Flash is also designed for software engineering and agent workflows, making it another candidate to compare.

The coding environment matters: access to files, a terminal and tests determines what the model can investigate and verify. Ask for evidence that the change works, and review the result before shipping it.

Try two models on one real task

Choose a recurring task with a known good outcome. Give two models the same information and instructions, then check:

  • Did they preserve the important facts and follow the constraints?
  • How much correction did each answer need?
  • How long did the whole job take, including your review?
  • What did each attempt cost, including retries?

Keep the lighter option if it meets your standard. Move up when the extra reasoning saves enough editing or repeated attempts to justify it. Revisit that choice as your tasks and the models change.

Use different models within the same work

Gemini 3.8 Flash, Claude Haiku 5.5, Claude Opus 5.5, GPT-6.1 Sol and GPT-6 Luna are all available in Drephos. You can choose a different model for the next message while keeping the same shared conversation and its history.

That gives a team a practical way to experiment: extract the facts with one model, develop the recommendation with another, then let a colleague question it in the same chat. A shared project keeps related conversations together when the work grows beyond one thread.