Custom data collection for physical AI

Tell us the data you need.
We have 130,000 professionals in the field.

You set the task, environment, viewpoint, volume and quality bar. We handle sites, operators, rigs, consent, redaction and QC.

2 to 6 synchronized cameras on live worksites, calibrated per unit. Enough views to recover 3D hand pose and 6-DoF trajectories, not just RGB.

Built on Miso's service network
15M+
Completed bookings
130K+
Service providers
1.5M+
Locations served
Residential and commercial customers in Korea, the UAE and the United States
Section 01  /  Capability

Volume, diversity, and rigs already in the field.

Every team that evaluates us pressure-tests the same three things. Here is where we stand on each.

01 / Volume
Scale comes from work already booked
We don't build sets or hire crews to reach a number. 500,000 hours a month of live job flow, already scheduled and already paid for. Scaling a collection means widening a filter, not recruiting.
02 / Diversity
Concentration caps, not just variety
Diversity you can audit, not hope for. We record site, operator and environment on every clip, so you can cap how much of a delivery comes from any one of them. The caps go in the written acceptance bar and we reject against them.
03 / Rig operations
Hardware we already run
Multi-camera means shipping rigs, training operators and proving calibration held. We run that loop today: rigs built and calibrated per unit, operators trained per collection, calibration verified before data ships.

Collections running now in Korea, the UAE and the United States. Everything above is a loop we are already operating, not a capability we would stand up for you.

Section 02  /  Coverage

Six environment classes. Three countries.

Your robot will not spend its life in one kind of room. Pick the settings you need and we line up sites and operators in each of them.

Env 01
Residential
Occupied homes. Kitchens, bathrooms, living and storage space, with whatever clutter and layout the household already has.
Env 02
Commercial
Offices, hotels and service venues. Shared space, turnover between guests, and front-of-house work done against the clock.
Env 03
Retail
Storefronts and back rooms. Restocking, facing, packing and inventory work, done around customers and a fixed planogram.
Env 04
Industrial
Plants, utility rooms and equipment bays. Machine handling, inspection, maintenance and heavy tool work.
Env 05
Manufacturing
Production and assembly floors. Repetitive precision work, part handling, fixtures and line-side tasks.
Env 06
Construction
Unstructured sites. Install and fit-out work, heavy tool use, and conditions that change between one visit and the next.
Somewhere else
Need a setting we have not listed? Tell us what it is and we will line up sites and operators for it when we scope the collection.

Collection runs in Korea, the UAE and the United States under one agreement. Running the same task family in three regions is the cheapest way to find out whether a policy transfers or whether it learned one country.

Section 03  /  Capture

One camera to six, synchronized.

Running six synchronized cameras in an occupied worksite is a different operational problem than running one, and it is the range we already work in. You pick the count when we scope the collection.

1 camera
Monocular

One RGB sensor, light enough that operators stop noticing it. That is why this is the setup we can run for full shifts across a lot of sites.

  • Single RGB sensor
  • Egocentric or fixed mount
  • Runs for full shifts
2 cameras
Stereo

A calibrated pair on a fixed baseline. Use it when you need disparity, depth or 3D structure around the hands and the object.

  • Synchronized camera pair
  • Known baseline & intrinsics
  • Depth-derivable output
3 to 6 cameras
Multi-view array

A synchronized array covering the work volume from several angles. Enough overlap to triangulate 3D hand pose and 6-DoF object trajectories, and to keep the hands visible when one view gets occluded.

  • Frame-level sync across the array
  • Full intrinsics & extrinsics per unit
  • 3D hand pose & 6-DoF derivable
  • Occlusion-resistant coverage

Camera count, array geometry, calibration, sync, resolution and frame rate are fixed per collection and documented with the data. Tell us what your model reads and we will confirm the setup before anyone starts recording. Delivery format is part of the spec. Name the schema you work in, LeRobot, RLDS or your own, and the export gets built into the collection.

Section 04  /  Operations

How a collection runs.

Rigs, training and quality control are where a collection breaks. We run the whole loop in house, so a multi-camera collection does not add a vendor to your chain or a handoff to your timeline. The acceptance bar is agreed in writing before collection starts, and every step below runs against it.

01 / Rig build
Rigs assembled and calibrated before they leave us, with intrinsics and extrinsics recorded per unit across the whole array. Units go out to the operators on your collection and stay tracked by serial for the whole run.
02 / Operator training
Each collection gets its own operator cohort, trained on mount, framing, array placement, and when to start and stop for that task.
03 / On-site capture
Operators record while doing the paid job, so capture rides a schedule that already exists. Operator and site owner consent is taken when the job is booked.
04 / Calibration check
Returned data is checked for frame sync across the array, baseline and calibration drift. Clips that fail get rejected, not patched.
05 / Concentration audit
Delivered hours are checked against the site, operator and environment caps in your acceptance bar before the batch closes.
06 / Delivery
Faces and PII are removed before footage leaves Miso. Video ships with a per-clip record carrying site, operator, environment and calibration, plus task segmentation when the collection calls for it.
Section 05  /  Data

What we capture.

Deformable manipulation: textiles, garment handling, stain treatment, alterations and sewing. Cloth, fine motor tool use, and outcomes you can grade.
Articulated objects, long horizon multi-step work, and the parts where it goes wrong and gets fixed.
Wiping a kitchen counter and wiping down a machine housing are closer than they look. Collect the same task family in six settings and a model has something to generalize from. Collect it in one and it learns the room.
These are examples, not a catalogue. Describe the work you need and we scope it.

Raw video with a per-clip record is the baseline. Task segmentation is available, and interaction labels, outcome states and any other fields get defined with you before collection starts.

Section 06  /  Execution

Why we turn a spec around faster than a bespoke shoot.

A bespoke collection starts by finding sites, recruiting people and getting them scheduled. For us all three already exist. Miso runs a live service network across Korea, the UAE and the United States, and Miso Motion adds capture, consent, redaction and quality control on top of work that is already booked.

  • No recruiting, no set building, no location scouting. Your start date is a scheduling question, not a construction project.
  • With 130K+ providers we can hold your spec constant and vary the site, or hold the site constant and vary the task. Either direction is a different filter on the same network.
  • Diversity is a number in your acceptance bar, not a hope. Site, operator and environment are on every clip, so concentration caps are auditable at delivery.
  • Rig builds, operator training and calibration checks all run in house. Standing up a synchronized 2 to 6 camera array on an occupied site adds no vendor and no handoff.
  • Every clip is a professional doing paid work, not an actor running a script. Technique stays consistent across operators, which is what imitation learning needs.
  • Three countries under one agreement, so cross-region coverage is a line item instead of three procurement cycles.
  • Licensing is structured per buyer: exclusive, non-exclusive or perpetual, agreed with the collection.
Get Started

Send us the spec.

Four fields and we send real clips with the per-clip records and capture specs behind them.