AI Engineer - Remote | Hybrid

IT - Software / DB / QA / Web / Graphics / GIS

About the Employer

Job Description

Join Our Team

Allion Technologies Sri Lanka delivers innovative, scalable software solutions for businesses around the world – combining global expertise, agile development, and a client-first approach to drive digital transformation and long-term growth across industries.

AI Engineer

Remote / Hybrid

Responsibilities

  • Design and implement RAG pipelines: document ingestion, chunking strategies, vector store management, retrieval evaluation, and response generation
  • Build and maintain AI agent systems using agent harness frameworks (LangChain, LlamaIndex, AutoGen, Google ADK or crewAI)
  • Integrate with LLM APIs (OpenAI, Anthrop ic, Mistral, Bedrock) and manage prompt lifecycle, versioning, and evaluation
  • Engineer and iterate on prompts system prompts, few-shot examples, chain-of-thought patterns, and structured output formatting
  • Implement and manage agent memory short-term context windows, long-term persistent memory, episodic and semantic memory stores
  • Build and maintain knowledge bases, vector databases (Pinecone, Weaviate, pgvector, Chroma), document pipelines, and hybrid search systems
  • Evaluate and benchmark model outputs — build evals, track regressions, and measure RAG retrieval quality (RAGs or similar)
  • Collaborate with ML engineers on fine-tuning small language models (LoRa, QLoRA, SFT on open models)
  • Write clean, testable Python FastAPI, Fast MCP services, async workers, and data pipelines

Requirements & Skills

  • 3–4 years of software engineering experience with a minimum of 2 years working on AI/ML systems in production
  • Strong Python async programming, type hints, testing (pytest), packaging
  • Hands-on experience building RAG systems end-to-end including chunking, embedding, retrieval, and re-ranking
  • Experience with at least one agent framework LangChain, LlamaIndex, CrewAI, AutoGen, or equivalent
  • Working knowledge of LLMs context windows, tokenisation, temperature/sampling, model selection trade-offs
  • Practical prompt engineering structured outputs, few-shot prompting, chain-of-thought, tool-use prompting
  • Experience with vector databases Pinecone, Weaviate, pgvector, Chroma, or similar
  • Familiarity with agent memory patterns conversation history management, summarisation, long-term memory persistence
  • Comfortable working with REST APIs, cloud services (AWS/Azure), and containerised deployments (Docker)
  • Experience fine-tuning small language models (Mistral 7B, Phi-3, LLaMA 3) using LoRA/QLoRA
  • Knowledge of reinforcement learning from human feedback (RLHF) or preference optimisation (DPO, PPO)
  • Familiarity with evaluation frameworks RAGAS, TruLens, LangSmith, or custom evals
  • Experience with knowledge graph construction or ontology-based retrieval
  • Exposure to multi-agent orchestration patterns

Apply now: [email protected]

Let's grow together with Allion! www.alliontechnologies.com