Senior AI Engineer

At ZharfaTech, we’re pioneers in building AI-driven solutions that redefine industries. From generative AI to real-time inference systems, our team tackles cutting-edge challenges using the latest tools and frameworks.

Position Overview

We’re seeking a Senior AI Engineer with deep expertise in modern AI/ML frameworks, scalable deployment, and end-to-end system design. You’ll architect and optimize AI solutions using tools like PyTorch, Hugging Face Transformers, and vector databases, while ensuring robust, secure, and efficient deployment.

Key Responsibilities

  • AI/ML Development:

– Design and train state-of-the-art models (NLP, computer vision, generative AI) using PyTorch, Hugging Face Transformers, TensorRT-LLM, and MLX.

– Optimize inference with vLLM, AutoGPTQ, AutoAWQ, and FastEmbed for efficiency.

– Implement techniques like quantization (unsloth) and distributed training for performance scaling.

  • API & Backend Engineering:

– Build high-performance REST APIs using FastAPI and document endpoints with Swagger/OpenAPI.

– Secure APIs with TLS/SSL and integrate authentication/authorization workflows.

– Design MCP (Model Compute Platform) server architectures for scalable AI workloads.

  • Deployment & Infrastructure:

– Containerize AI systems with Docker and orchestrate multi-service environments using Docker Compose.

– Deploy models on cloud platforms (AWS/GCP/Azure) with CI/CD pipelines (GitHub Actions, Jenkins).

– Manage vector databases (Qdrant, pgvector) and relational databases (PostgreSQL) for AI applications.

  • Collaboration & Innovation:

– Experiment with emerging tools like CrewAI (multi-agent frameworks) and OpenHands (gesture recognition).

– Mentor junior engineers and lead cross-functional teams to align AI solutions with business goals.

Required Skills

  • Core AI/ML:

– ۵+ years of hands-on experience with Python, PyTorch, and Hugging Face Transformers.

– Expertise in model optimization (quantization, pruning) using TensorRT-LLM, AutoAWQ, or unsloth.

– Familiarity with generative AI workflows (LLM fine-tuning, RAG architectures).

  • Deployment & DevOps:

– Proficiency in Docker, CI/CD pipelines, and cloud platforms.

– Experience with vector databases (Qdrant, pgvector) and PostgreSQL.

– Knowledge of TLS/SSL and API security best practices.

  • Tools & Frameworks:

– FastAPI, Swagger, REST API design.

– Big data tools (Spark, Dask) and workflow orchestration (Airflow, Prefect).

– MLX (Apple Silicon optimization) and FastWhisper (ASR systems) is a plus.