Computer vision
Models that detect, classify and measure what's in images and video, in the cloud or on the device.
Table 1. At this threshold.
| Detections | Precision | Recall |
|---|---|---|
| 6 | 0.83 | 0.83 |
What we build
Computer vision solutions
We build models that inspect, count, read and monitor from your cameras and images, running in the cloud or on devices at the edge.
Visual quality inspection
Detect defects on production lines against agreed acceptance criteria.
Detection and counting
Count and locate items, vehicles or stock on shelves from camera images.
Document images
Read text, tables and fields from scanned documents and photos (OCROptical character recognition: reading printed or handwritten text from images.).
Safety monitoring
Detect people in restricted zones or missing personal protective equipment in video.
Damage assessment
Assess damage to vehicles, property or goods from photos to speed up claims and returns.
Grading and sorting
Grade produce, parts or materials by size, colour or quality from line cameras.

How to get started
Three steps to a plan for computer vision
- 1Tell us what you needUse the project form or book a call. A few sentences about the goal is enough to start.
- 2Free technical consultationWe go through your goals, users, existing systems and constraints with you.
- 3Your planA detailed plan covering the right tech stack, architecture, timeline and budget. Then you decide.
How it works
Four kinds of model we build
The same camera frame, as each kind of model outputs it. Choose a carton to see its result and what the line does with it.
- It outputs
- A box around each object, with its class and confidence.
- Measured by
- mAP, using IoU to decide whether a box matches the true one.
Carton 2
- Class
- carton · 0.97
- Defect
- dent · 0.91
- Box
- [308, 138, 54, 46]
Diverted to inspection
Tuned to your line
We set the confidence threshold with you. Move it to see which detections count and what gets missed, and point at any box to see what it is.
Correct
7
False alarms
1
Missed
2
Precision
88%
Recall
78%
Glare, shadows and price labels can look like products; half-hidden products score low. The threshold is set from validation on your own images, against the cost of a false alarm and of a miss.
Selected detection
Glare on the shelf edge · 0.55
Counted: false alarm
Our approach
How the work runs
Define the task
Classification, detection or segmentationOutlining the exact pixels that belong to each object, rather than drawing a box around it., with acceptance criteria per class.
Collect and label
Gather representative images, including rare cases, and label them to written guidelines.
Train
Adapt pretrained models to your images through transfer learningStarting from a model already trained on millions of images and adapting it to yours with fewer examples..
Validate in real conditions
Test across lighting, angles and cameras, with precision and recallPrecision: how many detections were correct. Recall: how many real cases were found. per class.
Deploy and monitor
Run in the cloud or on edge devicesHardware on site, such as an industrial PC or smart camera, that runs the model next to where images are captured., and monitor performance as conditions change.
What we'll need from you
Having these ready keeps the work moving.
Sample images or camera access
Images or video from the real cameras and conditions, including rare cases.
Acceptance criteria
Quality or operations staff who define each class and review labels.
Site access
Help with camera placement, lighting and network or edge hardware on site.
Privacy sign-off
Review of what may be captured and stored, where images include people.
Who's on the project
Our team, working with quality team from yours.
1, 1, 1, 1, 2
1 Mayrian 2 Your organization
Computer vision engineer: Selects, trains and evaluates the vision models.
Services
Computer vision services
Computer vision consulting
Use-case selection, a camera and image assessment, and acceptance criteria for each class.
Data collection and labelling
Representative images, including rare cases, labelled to written guidelines.
Model development
Pretrained models adapted to your images and tested in real conditions.
Deployment, integration and support
Inference in the cloud or on edge devices, connected to your line systems or apps, and monitored after launch.
Deliverables
What you receive
Appendix A. What you receive
All code, data and documentation are handed over in your accounts and repositories.
In your hands
What the documentation looks like
Every project ends with documents your team can run with. Here is an excerpt of one of them.
Carton inspection: dent class
Line 2 cameras · 4,800 labelled images
Label as a dent when
- The surface is pushed in by 3 mm or more
- Draw the box tight to the deformed area, not the whole carton
- Creases along fold lines are not dents
Per-class results on the validation set
| Class | Precision | Recall | mAP@0.5 |
|---|---|---|---|
| Dent | 0.93 | 0.90 | 0.92 |
| Tear | 0.91 | 0.86 | 0.89 |
| Missing label | 0.98 | 0.97 | 0.98 |
Measuring success
How success is measured
What we report on in computer vision projects. Which measures apply, and their targets, are agreed with you at the start.
Table 2. What we report, and when. Targets are agreed with you at the start.
| Measure | Reported |
|---|---|
| Precision and recall per class | Validation set, per class |
| mAP | Validation set |
| IoU | Validation set |
| Character or field accuracy | Validation set, for OCR |
| False reject rate | Pilot on the line |
| Throughput and latency | On the target hardware |
Measure
Precision and recall per class
For each defect or object type, how many detections were right and how many real cases were caught.
Reported
Validation set, per class
How we score accuracy
In the evaluation report, a detection counts only if it overlaps the labelled box enough. Drag the predicted box to see the overlap change.
Drag the blue box, or focus it and use the arrow keys.
Intersection over union
0.35
Too little overlap: scored as a miss and a false alarm
Detection accuracy (mAP) is reported at an IoU threshold, commonly 0.5, or averaged from 0.5 to 0.95 for stricter tasks such as measuring defects.
Readiness check
Are you ready for
computer vision?
Five questions, about a minute. You'll see what to settle first and a sensible starting point.
Datasheet
Readiness for computer vision
Questions to answer about your data and organization before building
Q1. Can you describe exactly what the system should find, with acceptance criteria per class?
Q2. Do you have, or can you collect, images from the real cameras and conditions?
Q3. Are examples of rare cases, such as defects, available or collectable?
Q4. Do you know where the model should run: in the cloud, on a server on site or on the device?
Q5. If images include people, have privacy rules been reviewed?
How to start
From first call to production
Start where you are. Each step ends with a decision, so you commit to the next one only when it makes sense.
Protocol
How an engagement runs
Each step ends with a decision on whether to continue
Free technical consultation
One or two sessions
Talk through the goal, the data you have and the systems involved.
Outputs
- (a) A shortlist of use cases, ranked by value and feasibility
- (b) A recommended next step
- (c) A plan covering stack, architecture, timeline and budget
Estimate the value
What it could be worth to you
Enter your own figures. The formula is shown, and the estimate can go with your enquiry.
Estimate
Manual inspection time
Computer vision · from your own figures
Build or buy
When you don't need
a custom build
Part of the free technical consultation: when an existing product covers the need, we recommend it instead of a custom build. These are the options we weigh, alongside the tools you already have.
Related work
Existing products that may be enough
- [1]Google Cloud Vision, Amazon Rekognition or Azure AI VisionWhen general tasks such as common object labels or reading printed text.
- [2]Document AI services such as Azure AI Document Intelligence, Amazon Textract or Google Document AIWhen standard documents such as invoices, receipts and forms, using their prebuilt models.
Our approach
When a custom build is worth it
- Your products, defects or parts aren't in general models
- Accuracy must be proven per class under your conditions
- Inference must run on edge devices or on site
- Results must drive line systems or applications
Technologies and standards
Chosen for your project
Built on the cloud you already use. Choose yours to see the services involved; we recommend the full stack in the free technical consultation.
Table 2. Managed services for each layer, by cloud. The highlighted column is the one you chose.
| Layer | AWS | Azure | Google Cloud |
|---|---|---|---|
| Ready-made vision and OCR | Amazon Rekognition, Amazon Textract | Azure AI Vision, Document Intelligence | Cloud Vision API, Document AI |
| Custom model training | Amazon SageMaker AI | Azure Machine Learning | Gemini Enterprise Agent Platform |
| Labelling | SageMaker Ground Truth | Azure ML data labeling | CVAT or Label Studio on GKE |
| Edge devices | AWS IoT Greengrass | Azure IoT Edge | LiteRT on the device |
Also runs on any of the three: NVIDIA Jetson, ONNX Runtime, Ultralytics YOLO, OpenCV.
Models and libraries
- PyTorch
- TensorFlow
- OpenCV
- Ultralytics YOLO
- Detectron2
Labelling
- CVAT
- Label Studio
Edge deployment
- NVIDIA Jetson
- ONNX Runtime
- TensorRT
- OpenVINO
Cloud vision services
- Amazon Rekognition
- Azure AI Vision
- Google Cloud Vision API
Next section
Generative AI
Assistants grounded in your content, evaluated before and after launch.
Previous: Machine learningData Analytics & AI
6 Generative AI
Can it read, write and act for us?
Questions
Common questions
about computer vision
How many images do we need?
We establish it with a pilot on your own images. Starting from pretrained models keeps the number needed down, and the pilot shows the accuracy to expect before a full rollout.
Can it run on our existing cameras?
In many cases. We assess your camera feeds, resolution and placement early, and recommend changes where they limit what a model can see.
What about people's privacy?
We design systems that avoid identifying people where it isn't needed: faces are blurred, images can be processed on the device, and only the results are stored.
How is a project priced?
Well-defined scopes are delivered as fixed-price engagements; when requirements are still evolving, we provide a dedicated team instead. Either way, the free technical consultation ends with a plan covering tech stack, architecture, timeline and budget, so you know the cost before work starts.
