Anshuman Tiwari tanshuman145@gmail.com

AI systems engineer, India

I build the parts of AI systems that have to be correct.

Attention kernels that stay numerically exact without materialising an N×N matrix. An orchestration runtime where the model returns a plan and never an action, so the runtime can reject it before anything runs. RAG pipelines that hold up at 80,000 users.

Currently AI Engineer at Juristo AI Labs. B.Tech ECE (IoT) at MMMUT Gorakhpur, 2027.

  • 80,000+users on a platform I built
  • 40%LLM inference cost cut
  • O(N)attention memory, down from O(N²)
  • 2009LeetCode Knight rating

Experience

AI Engineer

Juristo AI Labs

Dec 2025 — present
  • Built and scaled a production Next.js AI platform spanning APIs, authentication, background jobs and frontend delivery for 80,000+ users.
  • Engineered a 4-stage RAG pipeline — document parsing, chunking, embeddings, retrieval — to deliver context-grounded legal responses.
  • Integrated LLM inference through OpenRouter behind a multi-agent routing layer that dispatches queries to the right tools and workflows.
  • Cut LLM inference cost 40% with prompt caching and selective routing of suitable workflows to lighter models.

Full Stack Developer Intern

Wecofy Pvt. Ltd. Letter of Recommendation

Jun 2025 — Aug 2025
  • Built a MERN commerce and delivery platform covering product listings, cart, order lifecycle and delivery operations.
  • Implemented Socket.io real-time tracking and notifications across three stakeholder groups: customers, vendors and delivery partners.

Selected work

Everything here is open source. Where there's a live link, it runs the real thing — not a video, not a mockup.

Jabrod

Agentic AI SaaS — founded and built

An agentic platform that automates business workflows end to end: document ingestion, vector storage and agent orchestration, serving 3,000+ users.

  • Next.js
  • Vector DB
  • Agents
  • RAG

Online Meeting Platform

Low-latency video with in-call AI

Real-time meetings over WebRTC with Clerk authentication and OpenAI APIs powering in-call assistance.

  • Next.js
  • WebRTC
  • Clerk
  • OpenAI APIs

College Resources Portal

Academic materials, notices and club management

An institutional platform with role-based access, built for and used by students at MMMUT.

  • Full stack
  • RBAC
  • MongoDB

mini-rag

A compact retrieval-augmented generation pipeline

Chunking, embeddings, retrieval and grounded answering, kept small enough to read in one sitting.

  • TypeScript
  • Embeddings
  • Retrieval

Skills

Languages
C++, C, Python, JavaScript, TypeScript, SQL
AI / ML systems
LLMs, AI agents, agent orchestration, RAG, tool calling, embeddings, vector databases, prompt caching, GPU kernel programming (Triton/CUDA), attention & transformer internals
Frameworks & APIs
PyTorch, Triton, LangChain, LangGraph, OpenAI APIs, OpenRouter
Web & backend
Next.js, React, Node.js, Express, Socket.io, REST APIs
Databases & cloud
MongoDB, PostgreSQL, Firebase, Supabase, Convex, GCP, AWS
Tools
Git, GitHub, Docker, Postman, pytest, GitHub Actions

Competitive programming

  • LeetCodeKnight · 2009
  • CodeforcesSpecialist · 1508
  • CodeChef4 star

Education

Madan Mohan Malaviya University of Technology, Gorakhpur

B.Tech, Electronics & Communication Engineering (Internet of Things) · CGPA 8.39

2023 — 2027

Udaya Public School

CBSE Class XII — 90% · CBSE Class X — 96.2%

2021 — 2023

Get in touch

I'm most interested in work where correctness is load-bearing — inference systems, agent runtimes, kernels. If that's what you're building, I'd like to hear about it.