Egocentric & POV Datasets

We collect net-new data, capture it to your requirements, and enrich or prepare datasets you already hold - across text, image, video, audio, and behavioral data.

Three ways to get started​

01 - See samples first

Start with sample clips.

Sample footage from existing egocentric collections. Check capture quality, framing, resolution, and metadata before scoping.

02 - Collect to spec

Net-new data, collected to your spec.

You define the environment, activity set, contributor profile, and device. We deploy and return footage structured to your schema.


03 - Extend coverage

Scale an existing dataset.

Extend an existing POV dataset into new settings, longer sequences, or additional markets. Tell us the gap, we scope the capture.

Datasets by environment

Robotics

Household video

First-person household video across kitchens, bedrooms, bathrooms, living areas, garages, and gardens. Real people complete cooking, cleaning, laundry, sewing, repair, tool use, gardening, and everyday object-handling tasks for embodied AI and imitation learning.


Robotics

Industrial video

First-person capture on factory floors, workshops, and production lines. Covers assembly, inspection, machine operation, tool handling, maintenance, repair, safety procedures, and repeatable step-based workflows for industrial robotics and computer vision.


Robotics

Warehouse video

First-person warehouse and logistics activity across receiving, storage, picking, packing, sorting, scanning, pallet handling, loading, and inventory checks. Built for robotic manipulation, workflow recognition, navigation, and human-object interaction models.


Computer vision

Commercial video

First-person activity across offices, hospitality venues, service environments, and other commercial spaces. Covers equipment setup, cleaning, stock handling, food preparation, maintenance, customer service, and repeatable workplace procedures.


Computer vision

Retail video

In-store activity from shopper and staff perspectives. Includes shelf browsing, product selection, basket handling, restocking, price checking, barcode scanning, checkout, returns, and customer interactions for retail computer vision and behaviour analysis.


Navigation

Outdoor navigation video

Street-level first-person movement through public and outdoor environments. Covers walking, wayfinding, road crossing, obstacle avoidance, cycling, public transport, and interaction with signs, pathways, and changing terrain.


Capture devices Smartphones, cameras – head-mounted, chest-mounted, or handheld
Resolution 1080p / 30fps standard; up to 4K / 60fps on capable devices
Format MP4 video with metadata in JSON or CSV
Audio Ambient sound and voice captured where the task requires it
Framing Hands and manipulated objects kept in frame
Annotations Per-clip scene, location, lighting, and motion metadata as standard. Activity and sub-step labels, on-frame object inventory, hand-object interaction events, and first-person narration on request.
Volume Scaled to your project – from pilot batch to ongoing collection
Licensing Project-based or exclusive. Rights and provenance defined upfront
Delivery Existing datasets shared on request; new collection scoped to your project, delivered in days to weeks depending on volume and complexity.
Pricing Priced per hour of data, based on industry, scope, and technical requirements. No seat licenses, no retainers.

Your data team has better things to do

Acquirox handles data collection, dataset preparation, labeling, QA, and delivery - so your engineers stay focused on model development.

FAQ

Real. Every clip is captured by a verified contributor performing a real action in a real environment. No staged studio footage and no synthetic generation - the variability comes from real people in real settings, which is what embodied models need to generalize.
Phone-based capture - head-mounted, chest-mounted, or handheld. Device, resolution, and framing are scoped to your spec before collection begins. Because the network records on consumer phones, capture stays within what contributors can reliably produce at scale.
First-person image and video across defined environments - household, industrial, warehouse, commercial, retail, and outdoor. Scenarios are scoped to what contributors can perform reliably on a phone, covering everyday actions, manipulation, and movement. Tell us the environment and action list and we will confirm feasibility.
18M+ verified contributors across 150+ countries, concentrated in Asia, Latin America, and Africa. This gives you environmental and demographic range across settings, and coverage is scoped to your target markets before a project starts.
Human in the loop on every submission. 2-3 contributors cross-check each clip against your spec, with automated integrity checks on format, resolution, and framing running first. Senior Acquirox reviewers set the golden standard on flagged and edge-case clips, so the dataset stays consistent as it scales.
Yes, for straightforward tasks - action labels, object tagging, and event marking on the clips we capture. This lets you order capture and annotation together and receive labeled footage in one delivery. Complex annotation is scoped case by case.
Defined upfront. Project-based or exclusive licensing, with documented provenance and contributor consent. You know what rights you hold before capture begins, and exclusive datasets are held only for your team.
MP4 / H.264 video with structured metadata in JSON or CSV. Per-clip metadata covers scene, environment, and action label as standard, so footage arrives organized and ready for your pipeline.
Priced per hour of footage. The rate depends on exclusivity, technical requirements, capture setup, and scenario complexity. No seat licenses, no retainers - we return a scope-based quote against your capture brief.
Samples in 48 hours. Full collection is scoped to your volume and scenario, delivered in days to weeks. We share sample clips from the pilot so you can validate the capture before committing to full volume.

© 2026 Acquirox. All rights reserved.