A small model should be assessed only against a clearly bounded business task and a representative sample of your own work. The decision does not start with which model is more popular. It starts with whether the system can produce an acceptable output within the required time, on the available device or with the available connection. For ERP and operational processes, one further rule applies: a model may prepare a suggestion, while an authorised person remains responsible for the official transaction.
A narrow task has a recognisable input, a limited output, and a person who can verify the result. Examples can include sorting short enquiries into predefined categories, extracting fields from a standardised document, or drafting a response from an approved knowledge base.
That type of task differs from open-ended advice, interpretation of unclear rules, or autonomous posting. The greater the business consequence, the more important input control, output review, an exception rule, and a decision record become.
Apple documents Foundation Models for app development. That documentation may be relevant when an implementation inside an application is under consideration. Supported devices, languages, availability conditions, and actual behaviour for a specific use case need verification before a project decision. General documentation alone should not be used to infer suitability for an individual business process.
A useful pilot does not begin with a broad request such as "introduce AI." It begins with one work decision.
Describe the task in one sentence:
Here is a hypothetical example: an accounting team wants to extract a suggested category from the description of an incoming document for review. The model does not post the document and does not change the master data. An operator reviews the suggestion, selects the final category, and performs the official ERP transaction. This arrangement makes usefulness measurable without transferring responsibility to the model.
When a task has no single acceptable output, separate the cases first. Domestic and foreign documents may need different rules, incomplete records may require a different procedure, and exceptions may need their own route. Business process screening can establish where a decision starts and ends before technology selection begins.
A small-model assessment for a business task requires real alternatives to be compared on the same sample. The alternative may be a manual procedure, an existing rule, another application workflow, or a model. A comparison without the same task and the same acceptance criterion does not support a useful decision.
Tie quality to the business output. For classification, it may be agreement with a confirmed label. For field extraction, it may be value accuracy and completeness of required fields. For a drafted response, it may be adherence to approved sources and the absence of unsupported claims.
Include the following in the sample:
Do not reduce the outcome to a single average number. Record the error type. A wrong classification, a missing required field, persuasive but unsupported text, and an unrecognised exception have different operational consequences.
Latency is not only the time required to generate an output. It includes input preparation, request transmission where applicable, waiting for a response, operator review, correction, and handoff to the next step. A short response without the required context can increase total work time.
Set the point at which measurement starts and the point at which it ends. For example, measurement may start when a work item is opened and end when a suggestion is ready for review. For a process where a user is waiting, consistency of duration matters as well as the average.
The question of working on a device or with a connection is not a technical afterthought. It affects availability, the flow of data, and behaviour when work is interrupted.
For on-device processing, check whether the specific device can perform the required task within an acceptable time and how the application behaves when resources are insufficient. For processing with a connection, check what happens with a weak or interrupted connection, how the user receives processing status, and whether there is a clear return to the manual procedure.
Do not assume that either approach is inherently more suitable. The answer depends on the task, devices in use, connection conditions, input sensitivity, and the role of people in controlling the output.
Local language model evaluation means testing a model on a sample that represents your own process, with controlled access to data and a clear record of the expected outcome. In this context, "local" does not describe only where a model runs. It also describes an evaluation rooted in real business records and process rules.
Assign owners before testing:
A sample does not need to be large to be useful, but it must cover the actual range of work. Separate the set used to prepare instructions and rules from the set used for final checking. Otherwise, the process can be adapted to the very records later used as evidence of success.
For every record, retain the input in its approved form, the expected outcome or permitted range of outcomes, the system output, the reviewer decision, and an error-type label. Where input contains sensitive data, the data owner should define the minimum necessary sample content and the permitted processing method.
A ranking does not answer whether a particular operating approach fits your task. A simple decision framework is more useful.
Proceed to limited use when the sample shows acceptable quality against a pre-approved criterion, the duration fits the real workflow, exceptions have an assigned route, and the user understands when an output cannot be used without review.
Return to task design when errors recur in the same input group, when the reference outcome is not sufficiently clear, or when an operator spends more time checking than the procedure saves.
Retain the manual or existing procedure when quality, duration, or operating conditions do not support the business purpose. That finding is a valuable assessment outcome. It prevents an additional work layer with no clear benefit.
The ORKA approach separates responsibilities: ORKA manages day-to-day collaboration and the process framework, ERP remains the location of official transactions, and Trueforce covers specialised engineering work when needed. This division helps preserve a clear boundary between a system suggestion and a business decision.
A small model may produce a useful result on a standard task and still be unsuitable for exceptions. A pilot should not hide difficult cases or treat them as success without review.
Quality may vary by language, document format, input length, and ambiguous terms. Feature availability may depend on the device, application settings, connection, or other conditions that need project-specific verification. A change to instructions, an application, or a workflow can alter results, so relevant checks should be repeated after a material change.
Responsibility remains the key limitation. A model does not know business context beyond the input provided and does not take responsibility for the consequence of its output. The process needs a person or rule that stops, supplements, or rejects an unreliable result.
Record the following before deciding on limited implementation:
If the process does not yet have a clear input, decision owner, or exception route, organise the process first. Then talk to the ORKA team about screening and evaluation on your own sample, without bypassing the responsibilities required by ERP and operational work.