Project Cursa - Robot Manipulation Video Annotator
About this role
About the Role
We are looking for detail-oriented annotators to help label robot manipulation videos for AI training purposes. You'll watch short videos of robots performing manipulation tasks (filmed from three synchronized camera angles) and produce precise, structured, natural-language descriptions of the actions taking place. This work directly supports the development of robotics AI models and requires strong written English, sharp observational skills, and the discipline to follow a detailed style guide consistently.
What You'll Do
-
Watch short robot manipulation videos, each filmed from three synchronized camera views (an overhead view and views from each of the robot's two wrist-mounted cameras).
-
Break each video into time segments and write clear, natural-language descriptions for each segment.
-
Apply labels at three levels of detail for each applicable segment:
-
-
Atomic motion (a few seconds) — a single small movement (e.g., "close fingers around the red handle")
-
Skill / subtask (several seconds to ~20 seconds) — a complete, meaningful action (e.g., "pick up the red block by its edge")
-
Task / goal (up to ~1 minute) — the overall purpose of a sequence of skills (e.g., "place all blocks in the container")
-
-
Ensure every moment of video is covered by a label at two or more of these levels — no gaps, including idle or pause moments.
-
Accurately describe exactly what happens, including when something doesn't go as planned (a dropped object, a failed grasp, a slipped grip). Precision matters more than making the robot look successful.
-
Cross-reference all three camera angles: use the overhead view to understand the overall scene and object identity, and the close-up wrist views to confirm exact contact and grasp details.
-
Follow a detailed style guide covering vocabulary for actions, spatial relationships, object descriptions, and manner of movement, applying it consistently across many episodes.
-
Participate in periodic calibration sessions to align your labeling with the team and the client's reference examples.
What We're Looking For
Required:
-
Strong written English — you'll write dozens of short, precise descriptive sentences per video and need to vary your language rather than repeating the same phrases.
-
Sharp attention to detail — able to distinguish small differences (a successful grasp vs. a fumble, a push vs. a drag, which specific object part is being touched).
-
Comfort following a detailed, structured style guide and applying it consistently, even in ambiguous or edge-case scenarios.
-
Basic comfort with spatial/mechanical description (left/right, above/below, naming object parts like handles, lids, or edges).
-
Reliable, self-directed work habits — this is often heads-down work with periodic check-ins rather than close supervision.
Nice to Have:
-
Prior experience with video annotation, data labeling, transcription, or QA work.
-
Familiarity with robotics terminology (grippers, end-effectors, manipulation) — helpful but not necessary, as the style guide is self-contained.
-
Experience with annotation tools such as Label Studio.
Why Join Welo Data?
✨ Limitless Flexibility
Project-based opportunities that fit your availability. Choose when and how much you want to contribute—fully remote, with complete autonomy.
🌱 Limitless Growth
Optional access to AI and Large Language Model workshops designed specifically for professionals like you. No coding required—just your expertise.
🌍 Limitless Support
Be part of a global contributor community with responsive guidance and support.
💡 Real Impact
Apply your expertise in the Legal field to influence the AI systems shaping the future of your industry—while collaborating with data professionals and expanding your skills.
How to Apply?
Apply now by answering a few quick questions to join our database and become part of our growing community.
About Welo Data
Welo Data, part of Welocalize, is a global AI data company with 500,000+ contributors delivering high-quality, ethical data to train the world’s most advanced AI systems. We’re building smarter, more human AI with a diverse community in 100+ countries.
At Welo Data, Limitless AI. Limitless You. isn’t just a slogan—it’s our promise. We build smarter AI through the power of human contribution, offering limitless opportunities for our global community to grow, contribute, and work on their terms.
Company at a glance
Welocalize, a Welo Global brand, serves localization teams through AI-enabled multilingual content solutions that enable enterprises to operate and scale globally. Welocalize combines AI, automation, and human expertise to support enterprises in more than 300 languages, enabling accurate, culturally aligned, and compliant multilingual content at scale. Welocalize’s Opal Platform comprises patented technology designed to automate multilingual content and improve workflow performance for enterprises. Opal functions as an agentic system that orchestrates AI, automation, and human expertise to coordinate content workflows across the lifecycle, improving speed, scalability, and operational control for global enterprises. These solutions operate within a secure and compliant environment supported by seven ISO certifications. Welocalize is headquartered in New York with offices worldwide.
Top Benefits
- Flexible schedule
- Remote work
- AI and Large Language Model workshops
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?