Multilingual Data Contributors: PDF Collection for AI Training
About this role
What We're Researching
We're running a paid study on multilingual document sourcing to improve AI text recognition and generation. High-quality, legally usable PDFs in various languages are essential for training robust machine learning models. Your contributions will directly support the development of better language processing tools.
How It Works
You will work asynchronously to find and submit public, legally usable PDF documents in your designated language. During this process, you will verify that each document meets our quality and licensing requirements. You will upload the files through our secure platform and provide basic metadata for each submission. We will review your uploaded documents to ensure they match the project guidelines before approving the task.
Who This Is For
We are hiring fluent readers of Telugu, Odia, Gujarati, Malayalam, Japanese, and Korean who know how to source public documents online. Ideal candidates are detail-oriented individuals comfortable navigating digital archives, public records, or open-source repositories. We welcome data annotators, researchers, and general language contributors who understand basic copyright and licensing rules.
What You'll Do
Source legally usable, public PDF documents in your designated language
Verify that each document meets open-source or public domain licensing requirements
Upload the collected files to our research platform
Provide basic descriptive information for each submitted document
Who Should Apply
Fluent reading comprehension in Telugu, Odia, Gujarati, Malayalam, Japanese, or Korean
Comfortable searching for and downloading digital documents online
Basic understanding of public domain or open-source licensing
Access to a reliable computer and internet connection for uploading files
Compensation
$150 per task
Ready to participate?
About Terac
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Learn more at terac.com or on YouTube at @jointerac.
Company at a glance
Terac is an AI-native two-sided marketplace headquartered in downtown San Francisco that connects enterprises with specialized expertise to high-quality professionals on demand. The company's mission is to fundamentally change how labor is allocated globally by using AI-driven matchmaking to connect experts with work opportunities based on thousands of signals about capability, performance, preferences, and potential. Currently placing thousands of experts per month across 50+ projects for 25+ clients, Terac serves leading AI research companies, market research firms, data companies training AI models, and research organizations. With a panel of 100,000+ professionals and ambitions to scale to 10 million, the company has achieved an eight-figure annual run rate and expects to exceed $10 million in revenue in the next year. Having raised $9 million in seed funding, Terac is preparing to launch a Series A round while rapidly expanding its team from 4-5 people to approximately 15 members, with plans to reach 30 by year end.
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?