Marisol Vega
Prompt Engineer · Seattle, WA · marisol.vega@example.com · linkedin.com/in/marisolvega
Summary
Prompt engineer with 4+ years shaping LLM behavior in production: eval suites, guardrails, versioned prompt pipelines, and RAG context design across 3 customer-facing assistants. Cut hallucination rate 58% and harmful-output escapes to zero via evals-gated releases and quarterly red-teaming. Working Python; fluent with LLM APIs (OpenAI, Anthropic) and eval frameworks (promptfoo, Braintrust).
Professional Experience
Prompt Engineer · Halide AI
Feb 2024 – Present
- Cut hallucination rate 58% across 3 production assistants by building a 400-case golden-answer eval suite and a versioned prompt pipeline with regression CI gating every change.
- Took harmful-output escapes to zero over 6 months of production by layering guardrails (input classifiers, output filters, refusal handling) and running quarterly red-team sprints against a 120-case adversarial set.
- Raised support-assistant answer accuracy from 81% to 93% by re-engineering RAG context: retrieval-aware chunking, citation-forcing prompt structure, and per-query context-window budgeting.
- Cut prompt-change turnaround from ~2 weeks to same-day by moving prompts out of application code into a versioned registry with diff review, canary rollout, and automatic eval runs on every merge.
- Halved LLM-judge disagreement with human raters (34% to 16%) by rewriting judge rubrics with anchored score descriptions and calibrating against 500 human-labeled transcripts.
Conversation Designer · Meridian Assurance
Sep 2021 – Jan 2024
- Lifted chatbot containment from 34% to 57% (~$700K/yr in deflected call volume) by rewriting 200+ intent flows, then leading the migration from intent trees to a retrieval-grounded LLM assistant.
- Cut escalation misroutes 41% by authoring the assistant persona and refusal-policy guide and scoring 1,200 transcripts/quarter against a rubric I designed.
- Reduced legal-review cycles from 3 weeks to 4 days by building a claims-language style guide and templated response library adopted across 5 support teams.
Projects
Open-source: judgecal
- Built a calibration toolkit for LLM-as-judge rubrics (agreement metrics, anchored-rubric templates, drift alerts) used by 2 companies in production and cited in an eval-engineering newsletter with 20K subscribers.
Technical Skills
- Prompt & eval engineering: prompt design (few-shot, chain-of-thought, structured output), LLM evals (golden sets, LLM-as-judge), guardrails, red-teaming, RAG context engineering, prompt versioning & regression CI
- Tools & platforms: LLM APIs (OpenAI, Anthropic), eval frameworks (promptfoo, Braintrust), Python, vector search (pgvector), Git & CI (GitHub Actions)
Certifications & Education
- B.A. Linguistics, University of Washington — Jun 2019
- DeepLearning.AI, Generative AI with Large Language Models — Mar 2023