AI Evaluation Engineer at SiteGround
Posted on
Sofia, Bulgaria
Hybrid
Industry: Internet
Job Type: Full-time
Experience Level: Mid-Senior Level
Job Description
Prompt engineering is a central part of this role on SiteGround’s AI products, including Coderick AI, AI Studio, and the WordPress AI Agent. You design, test, and iterate on prompts and agent behaviour, and make LLM-powered features measurable and production-ready through evaluation and observability across OpenAI, Google Gemini, and Anthropic.
Responsibilities
- Build evaluation datasets with ground truth and define task-specific metrics such as correctness, task completion, efficiency, cost, token usage, and safety.
- Run model bake-offs and before/after prompt experiments so every meaningful prompt change is measured.
- Design evaluations for agentic workflows, tool-calling correctness, and structured-output validity where there is no clean ground truth.
- Own production observability through trace analysis, per-tool dashboards, error classification, and session-level review.
- Root-cause failures and turn recurring patterns into prompt improvements and engineering tickets.
- Design, version, test, and continuously improve system prompts for domain-specific agents.
- Architect agent behaviour, including tool orchestration, multi-turn flows, and safeguards such as plan-before-execute, read-before-write, and always-confirm for destructive actions.
- Build reusable, modular prompt components for safety, tone, context, and tool documentation, and select models based on quality, latency, and cost.
- Work across the API layer agents act through, integrate REST endpoints, and help design guardrails that prevent agents from corrupting user data.
Requirements
- Strong prompt engineering and LLM application experience, with agents or LLM-powered features shipped to production.
- Hands-on evaluation experience: you have built evaluation datasets and metrics and measure results before claiming something works.
- Python experience and comfort consuming and integrating REST APIs.
- Familiarity with function and tool calling, RAG, agentic workflows, observability platforms, and current model families.
- Healthy scepticism toward AI output and a strong instinct for safety, failure modes, and guardrail design.
- Advantage: experiment design and statistics, including A/B testing, significance, and sampling.
- Advantage: LLM-as-a-judge evaluations calibrated against human labels.
- Advantage: red-teaming and adversarial testing, including prompt injection, jailbreaks, and structured-output abuse.
- Advantage: SQL, pandas, DuckDB, notebooks, or similar tools for exploring evaluation results.
- Advantage: Langfuse, Langchain, n8n, or Google ADK.
Benefits
- Competitive remuneration and a performance-based bonus.
- Shorter workday on Fridays, finishing at 3:00 PM.
- Extra paid days off with tenure, plus two extra paid days off for volunteering.
- Premium additional health insurance and annual medical check-ups.
- Modern offices across Bulgaria, plus flexible work options.
- Chef-prepared breakfast and lunch at HQ, free parking, and a metro shuttle.
- In-office gym, on-site sports, company-covered Multisport or Coolfit cards, and free massages.
- Training, conferences, knowledge-sharing meetups, and team events.