OUR SORTS & CASES
A job case open flat on the research bench
Every compartment holds a single sort, a single discipline of linguistic engineering. Read the six columns the way a compositor reads a California job case: top to bottom, knowing where each sort earns its place and how the sorts act together when they are set in a line. Each column below carries a full description with links to a deeper page for those who want the technical detail of the work.
A full solution rarely travels along one sort alone. Applied models are only as good as their training material, training material is only as good as its annotation, and annotation is only as good as the framing definitions behind it. Our habit is therefore to treat the six services as one connected suite and to tell clients honestly when a task needs more than one line of the case crossed, so the assembly is booked and scoped together from the start instead of patched together late.
Applied Models
Custom natural language models built to answer specific business questions. We choose the architecture honestly, train on your own material, tune the settings and measure the output. A model earns its place only when it stays accurate, fast and defensible after deployment.
Every delivered model parts with a short report of the runs that produced it and the test route that proved it, so the next person on the job can read the record before they retrain. That bookkeeping turns a one off result into a settled asset the team owns.
DETAILS
Annotation Pipelines
Structured labelling that turns raw text and speech into clean training data. Rigorous guidelines, double review streams and a clear record of every disagreement keep labels consistent, auditable and ready for the model training that follows.
Behind the scaffolds sits a real review bench: annotated samples are checked by a second reader and the measure of agreement is reported rather than assumed, so a client knows the true quality of the corpus before a model is built on it.
DETAILS
Corpus Design
Source selection, sampling frames and balance planning for any collection of text. A well designed corpus is the difference between a fragile demo and a durable result. We document provenance so every future claim can trace its evidence.
When you already hold a body of documents, we audit those shelves too, reporting which parts are safe to use, which need cleaning and which should be retired, before you spend a single hour of modelling on unreliable ground.
DETAILS
Evaluation Frameworks
Honest measurement is the heart of the shop. We build test suites, scoring protocols and benchmark cards so every system is judged against results it can stand behind in a report or before a reviewer, not against a single cherry picked headline.
Benchmark cards record the composition of the test set and the version of every dependency, so two runs can be compared with confidence and a later system is measured against the same ruler as the one before it.
DETAILS
Domain Adaptation
Moving a model from the general text it was trained on into a precise field, from medicine to law to logistics. Careful adaptation studies and continued training keep language reliable where it does the real work of the business.
We run controlled before and after comparisons so the change can be attributed to the method rather than to luck on one test set, and we confirm adaptation does not quietly harm the general cases the firm still leans on day to day.
DETAILS
Research Partnerships
Long horizon collaborations for teams who need language analysis at depth. We share methods, publish openly where the brief allows and measure progress against agreed research questions rather than against the calendar alone.
Because partnerships are taken on in limited numbers each quarter, a serious long term programme is best opened as a sitting research slot rather than an array of one off orders, so good questions get the room they need to mature.
DETAILS