Seeking a Linguist to support the development of AI-powered products, including Large Language Models (LLMs), generative AI systems, and voice-enabled technologies. The ideal candidate will bring strong linguistic analysis skills, experience working with language data, and familiarity with AI/ML workflows.
This role will focus on designing and improving language components through data collection, annotation, synthetic data generation, evaluation, and quality analysis. The Linguist will collaborate closely with linguists, data operations specialists, and machine learning engineers to enhance model training, alignment, evaluation, and AI agent capabilities.
Location: Hybrid/Onsite – 3 days per week onsite at one of the following preferred locations:
- Sunnyvale, CA
- New York, NY
- Burlingame, CA
- Redmond, WA
Key Responsibilities
- Apply expertise in syntax, semantics, pragmatics, sociolinguistics, and corpus linguistics to support LLM and generative AI development.
- Partner with linguists, data operations teams, and ML engineers on language data collection, curation, annotation, and localization initiatives.
- Develop, refine, and maintain annotation frameworks, schemas, and guidelines for AI training datasets, including instruction tuning, preference labeling, and reinforcement learning from human feedback (RLHF).
- Perform dataset evaluation and quality assurance activities supporting model pre-training, fine-tuning, and alignment.
- Support the creation of scalable approaches for generating synthetic annotated data.
- Assist with AI model evaluation through prompt-based testing, linguistic analysis, adversarial testing, and error identification.
- Conduct experiments to measure annotation quality, consistency, and impact on downstream model performance.
- Analyze language data and provide insights to improve AI system accuracy, reliability, and user experience.
Qualifications
- Bachelor’s degree in Linguistics, Computational Linguistics, Computer Science, Speech Science, Language Technologies, or a related field.
- 1+ years of experience in linguistics, language technologies, NLP, AI/ML data operations, or a related discipline.
- Strong understanding of linguistic concepts including syntax, semantics, pragmatics, sociolinguistics, and corpus linguistics.
- Familiarity with Large Language Models (LLMs), including training data practices, evaluation methods, prompting, and fine-tuning workflows.
- Experience with LLM evaluation approaches such as human evaluation, automated metrics, and adversarial testing.
- Experience developing or working with semantic ontologies, taxonomies, intent/slot frameworks, or similar language structures.
- Experience using AI agents, chatbots, or conversational AI systems.
- Experience with data analysis tools and processes, including SQL, spreadsheets, R, Unix, or similar technologies.
- Experience working with multilingual speech and text data.
- Ability to thrive in a fast-paced, collaborative environment with changing priorities.
- Native or near-native fluency in English and at least one additional language.
- Master’s degree in Linguistics, Computational Linguistics, Language Technologies, or a related field.
- Familiarity with machine learning frameworks, NLP libraries, and tools such as Hugging Face, spaCy, NLTK, or PyTorch.
- Exposure to statistical language modeling, ML pipelines, or AI training data workflows.
- Strong organizational skills with excellent attention to detail.
- Experience supporting multilingual language data projects involving speech and text.
- Ability to collaborate effectively across technical and non-technical teams.