Applied AI Engineer
Job Description
Block Labs is seeking an Applied AI Engineer in Malta to help design and deliver governed, production-grade AI agents and the supporting data and agent platform behind business decisioning. The work spans AI agent engineering, machine learning and data science, and the dashboards and surfaces needed for auditability, safety, and operator control.
Responsibilities
Start by delivering a production SQL BI analyst agent that operates as a Slack-native assistant. The agent answers business questions using governed SQL executed against the analytical warehouse, with validated queries, sanity-checked outputs, and cited evidence for every reported number.
- Build and own analyst agents end to end, including natural-language to governed-SQL translation over the analytical warehouse.
- Support executive P&L questions, post daily health briefings, and provide data-backed root cause analysis for metric movement.
- Enforce safeguards through SQL and schema validation, result sanity checks, and cross-checks against canonical daily reporting views.
- Extend into customer-facing agents with intent triage and routing, retrieval-grounded responses over versioned knowledge bases, and strict abort-and-escalate fallbacks.
- Implement multi-turn conversational state machines, localized brand voice, and escalation logic integrated with helpdesk and CRM systems through webhook ingestion, session lifecycle management, intent metadata tagging, and automated escalation tickets with pre-packaged tool context.
- Create a risk-stratified tool layer between agents and back-office APIs, including read-only context-gathering tools, information-first validation, multi-turn confirmation workflows, and a per-tool switch for low-risk mutations to move from lead approval to autonomous execution as evidence accumulates.
- Harden systems against adversarial input using prompt-injection screening, confidence-threshold freezes for sensitive intents, silent security escalation paths, and defenses against tool misuse and data exfiltration.
- Build agents up the autonomy ladder using LangGraph, the Anthropic Agent SDK / Model Context Protocol (MCP), or comparable orchestration frameworks.
- Engineer the closed feedback loop, including corrections capture, proven-query and semantic memory, decision audit logging, and evaluation harnesses with regression suites to ensure new tools or intents introduce zero degradation.
- Produce and productionise models that power decision signals, including churn, lifetime value, and bonus-sensitivity models for engagement.
- Develop composite player risk scoring across identity, payment, gameplay, bonus, and network signals, including collusion, bot-play, and multi-accounting detection, plus anomaly detection for treasury and payments.
- Ship models as governed signals, using versioned, SLA’d contracts with the decision engine and agents, including freshness, drift, and calibration monitoring and automated retraining paths served across real-time (Kafka/MSK), near-real-time, and batch (ClickHouse) tiers.
- Own multi-vector withdrawal risk scoring with cited rationale and confidence, evidence-aware aggregation, and automatic re-scoring when late evidence arrives.
- Codify business rules with domain owners in an auditable way by translating policy into deterministic, configurable rules, simulating and backtesting changes, designing holdouts and control groups, and running deep-dive analyses that inform both agents and executives.
- Build supervisor and approval surfaces, including review queues with one-click action proposal cards, searchable session replay exposing prompts, model outputs, reasoning chains, and tool calls, and a structured grading module feeding evaluation and fine-tuning datasets.
- Design and ship dashboards for decision audit views, agent performance dashboards, risk review queues, and KPI views built with the BI team, moving dashboarding toward AI-assisted anomaly detection and explanation.
Work is structured so humans set objectives, budgets, and approval gates, while AI and ML models author, score, simulate, and optimise within them. Execution is deterministic inside approved boundaries with complete logging. The platform is designed so wrong decisions can be reconstructed and audited, with governance and security controls aligned to real customer and real money outcomes.
Requirements
- 4+ years of experience in software, data science, or machine learning engineering, including 1+ years building LLM-powered agents in production with tool use and function calling, structured outputs, retrieval and memory, and multi-step orchestration using frameworks such as LangGraph or the Anthropic Agent SDK.
- Experience shipping a production RAG system, including grounding, chunking, retrieval quality, hallucination control, and refusal behavior.
- Ability to treat customer-facing agents as an attack surface, with concrete defenses for prompt injection, tool-call abuse, and data leakage through model outputs.
- Ownership across the production ML lifecycle: feature engineering, training, serving, monitoring, and retraining, with fraud, risk, or abuse detection experience as a strong signal (including imbalanced classes, adversarial users, and cost-asymmetric decisions).
- Statistical rigor, including experiment design, holdouts and control groups, uplift measurement, and score calibration.
- Evaluation discipline for non-deterministic systems, including evaluation harnesses and regression suites to catch quality drift.
- Strong Python for production services, comfort in TypeScript for review and approval surfaces, and strong SQL skills on columnar analytical databases (ClickHouse preferred).
- Capability to bring outputs to stakeholder-ready surfaces such as review queues, approval interfaces, dashboards, and lightweight internal apps using a front-end framework or tools such as Streamlit.
- Experience designing systems where model outputs feed deterministic execution, with clean boundaries, LLM observability and tracing (Langfuse, LangSmith, or similar), and end-to-end ownership of what ships.
Technologies
- SQL, Python, TypeScript, ClickHouse, Kafka, MSK
- Slack, LangGraph, Anthropic Agent SDK
- Model Context Protocol (MCP), Langfuse, LangSmith, Streamlit
- RAG
Nice to Have
- Experience in iGaming or other high-trust, transaction-intensive environments emphasizing security, fraud prevention, auditability, traceability, data integrity, and robust operational controls.
- Helpdesk or CS-platform integration experience such as Intercom or Zendesk, including webhooks, conversation APIs, agent-assist, or automation.
- Exposure to blockchain or crypto-native transaction flows, including on-chain data, wallet clustering, or stablecoin settlement.
- Experience with constrained optimisation, bandits, or reinforcement learning under business constraints (budgets, caps, exclusion lists).
- Experience with rule engines or decision-management systems, and Slack app development.
- Event-driven and streaming experience, including Kafka or MSK consumers, idempotent processing, and failure handling.
How We Work
- Fully remote with asynchronous-first communication, with preferred EU timezone overlap.
- Small, high-autonomy Intelligence team within the Data function, reporting to the Head of Data and coordinating with AI, BI, and Infrastructure Teams, and for customer-facing agents, with the Head of CS and product squads whose platforms the agents serve.
- Architecture decisions are documented and debated, with participation in design reviews and ownership of domain decisions.
- Autonomy is earned by evidence, starting in propose-only, human-gated mode and progressing one level at a time with simulation and outcome data.
- Multi-tenant scale is a design requirement from day one, with solutions expected to absorb additional operators without extra engineering effort.
Company Overview: Block Labs
Block Labs is a technology studio operating at the intersection of Web3, Artificial Intelligence, and iGaming. They build high-scale, production-grade platforms for the next generation of digital products, led by senior engineers, product strategists, and builders who prioritise architecture and long-term engineering excellence.