Sr. AI/ML Engineer
Apply or review the details on the original posting.
Apply for this roleOriginal posting on LinkedIn
🌟 We're Hiring: Sr. AI/ML Engineer! 🌟
We are seeking aSenior AI/ML Engineer (MLOps)to design, deploy,monitor, and continuously improve large-scale AI and machine learning solutions in production. This role focuses on ensuring the reliability, performance, governance, and operational excellence of bothGenerative AIandtraditional ML systems, while collaborating closely with data scientists, engineers, platform teams, and business stakeholders.
Key Responsibilities
Design and implement monitoring, alerting, and observability solutions for ML and AI systems, including model performance, data drift, latency, and AI operational health.
Build andmaintainautomated evaluation, testing, and validation pipelines for ML models, GenAI workflows, prompts, and agentic systems.
Investigate and resolve production issues related to model behavior, AI applications, data quality, integrations, and RAG/retrieval pipelines.
Manage model, prompt, embedding, and vector index lifecycles, including versioning, rollout strategies, and evaluation gates.
Collaborate with platform and infrastructure teams tooptimizeAI deployments for scalability, reliability, and cost efficiency.
Establish feedback loops using user interactions, business metrics, UAT findings, and expert reviews to drive continuous improvement.
Ensure compliance with governance, security, privacy, and operational best practices.
Create operational documentation, incident reports, and performance benchmarks.
Communicate technical insights, risks, and system health to both technical and non-technical stakeholders.
Mentor junior team members through coaching, knowledge sharing, and code reviews.
General Background
Bachelor's degree in Computer Science, Data Science, Engineering, Information Technology, ora relatedfield.
6+ years of experience inMachine Learning, AI Engineering,MLOps, or related software engineering roles.
ExperiencesupportingproductionAI/ML systems in cloud environmentsand agile delivery teams.
Strong problem-solving, communication, stakeholder management, and mentoring capabilities; willingness toparticipatein on-call support whenrequired.
Mandatory Skills
Strong Pythonexpertiseincluding Pandas, NumPy, and Scikit-learn for model maintenance, automation, and production support.
Advanced SQL skillsfor data analysis, operational checks, troubleshooting, and data quality validation.
Hands-on experience building or supporting productionGenerative AI solutions, including RAG, AI agents,LlamaIndex,CrewAI, or similar orchestration frameworks.
Strong understanding ofMachine Learning fundamentals,model evaluation techniques, drift detection, bias monitoring, and model performance management.
Proven experience inMLOpsand model monitoring, using tools such asPrometheus, Grafana, or cloud-native monitoring platformsto track model performance, data quality, latency, and operational metrics.
Experience working withMicrosoft Azure AI/ML ecosystem, including services such as Azure Machine Learning, Databricks, or related cloud-based AI deployment platforms.
Practical experience working inAgile environments (Scrum/Kanban).
Nice-to-Have Skills
Experience with Apache Spark,Dask, or other distributed data processing frameworks.
Familiarity with model serving and deployment platforms such as Azure ML Inference, Triton, SageMaker Endpoints, and API frameworks likeFastAPIor Flask.
Knowledge of Docker and Kubernetes for scalable ML deployments.
Experience with orchestration and ML lifecycle tools such as Airflow, Kubeflow, orMLflow.
Exposure to Infrastructure as Code (IaC) tools such as Terraform.
Experience mentoring engineers and driving engineering best practices across teams.
Ready to shape the future with us? 🚀 Apply now and join our dynamic team!