- Design and run offline A/B experiments to design optimal evaluation pipelines for benchmark tests offered in Project Moonshot, including LLM-as jury.
- Contribute to the development and maintenance of benchmark datasets, particularly generating realistic test cases in the context of Singapore.
- Research emerging GenAI evaluation best practices and open-source tools — translate key insights to inform product roadmap, and turn research findings into actionable prototypes for engineering teams.
- Educate and engage the developer community to build awareness of the challenges in benchmark testing, and champion the solutions provided in our library.
[What we are looking for]
- Background in Statistics or Data Science, or a related technical field.
- 1-3 years’ experiences as a data scientist, with strong grounding in experimental design and analysis.
- Proficiency in Python/ R is required. Experience in AI testing is a plus.
- Deep intellectual capacity combined with a practical, user-centric mindset.
- Excellent communication skills— excel in storytelling with data, and able to simplify complex concepts for varied audiences.