Based in St. Petersburg, FL
Lukas Babaliauskas
Senior Product Engineer — full-stack TypeScript & applied AI.
I build customer-facing SaaS end to end — React and TypeScript up front, Node.js, PostgreSQL and AWS behind it — and the applied-AI tooling that proves it still works.
Overview
Product judgment. Technical ownership. Hands-on execution.
I'm a senior product engineer with 8+ years building customer-facing SaaS across TypeScript/React, Node.js, PostgreSQL and AWS — plus applied LLM systems.
At NinjaOne I lead and manage a small engineering team delivering user-facing features end to end. Independently, I built and shipped KidViq, a kindergarten-management platform used by hundreds of schools, and created EvalShift, an open-source LLM migration and regression-testing system for AI agents.
- 8+
- years shipping customer-facing SaaS
- 100s
- of kindergartens ran on KidViq in production
- 1.1
- EvalShift — stable, open source, Apache-2.0
- 2,200
- tests behind the EvalShift CLI, with a 90% coverage floor in CI
Work
Customer-facing SaaS, shipped end to end.
From React/Redux and serverless AWS to leading a team that owns features from design to production.
-
NinjaOne
Software & Platform Engineer
Leads a small engineering team shipping customer-facing SaaS.
- Lead and manage a small engineering team delivering customer-facing SaaS features across React/TypeScript frontends and Node.js/TypeScript backend services.
- Own technical design and execution for complex user-facing features, partnering with product, design, and UX to turn requirements into fast, intuitive, maintainable workflows.
- Design scalable REST APIs, microservices, and data flows for high-traffic SaaS; optimize rendering, state management, API latency, and database queries for responsive product experiences.
- Build reusable React/NPM component systems and provide technical direction through architecture decisions, code reviews, mentoring, and engineering standards.
- Operate production delivery end-to-end with automated testing, CI/CD, containerized AWS services, security checks, and Infrastructure as Code with AWS CDK.
-
Byrider
Software Engineer
Led frontend development and serverless AWS integrations.
- Led frontend development across React/Redux applications, set architecture/component standards, built Node.js APIs, and integrated AWS serverless workflows (S3, Lambda, API Gateway, DynamoDB).
-
Infosys
Full-Stack Developer
Built Node.js APIs and Vue.js interfaces.
- Built Node.js REST APIs and Vue.js interfaces with reusable components, authentication, validation, logging, testing, and linting for more reliable releases.
Education
- 2011B.Sc. Mathematics & Computer ScienceLithuanian University of Educational Sciences · Vilnius, LT
- 2018Full-Stack Web DevelopmentDevMountain · Provo, UT
Cartridge A — EvalShift
Open-source LLM evaluation & migration system · creator
Don't ship a model upgrade on vibes.
Switching LLMs is a behavior change with no diff to review. EvalShift replays real, captured agent behavior on the model you have and the model you want, scores both — structure, semantics, LLM-as-judge and tool calls — and returns a verdict you can defend.
Migration verdict
FAIL
main_chat · gemini-3.7-flash → gemini-3.5-flash-lite
16 captured examples replayed on both models · 5 of 7 budgets within policy
- Confirms actions it never performed Asked to book coffee, it says it's on the calendar — without calling any tool.
- Routes to the wrong tool A weather question gains a
display_infocall. - Searches instead of answering Asked for an opinion, it issues
search_weband returns no text.
- Cost
- −58.8%
- Latency p95
- −75.2%
- Tool divergence
- 25.0%ceiling 10%
- Equivalence
- 61.4%floor 75%
Do not migrate: a −58.8% cost saving does not cover 25% tool divergence.
How it works
Capture
One decorator — or a wrapped OpenAI, Anthropic or Google GenAI client — records model calls, tool calls and outputs to disk; every capture point declares how it redacts.
Compare
evalshift comparereplays the suite on both models, round by round for multi-turn agents, and scores every example.Decide
Paired t-test or Wilcoxon after a Shapiro–Wilk screen, Cohen’s d with 95% CIs, Benjamini–Hochberg across every prompt × evaluator × slice. Budgets in
migration_policyturn it into pass, conditional or fail.Gate
The GitHub Action evaluates every suite on each pull request, posts one comment and holds a single
evalshift gatecheck.
What I built
- Built a local-first LLM migration and regression-testing system for AI agents that replays golden suites against candidate models and evaluates structural, semantic, LLM-as-judge, and tool-call behavior.
- Designed agent-specific regression checks for tool selection, arguments, and sequencing, plus statistical analysis with paired tests, Cohen’s d, 95% confidence intervals, and Benjamini-Hochberg correction.
- Added regression-budget verdicts, single-file HTML reports, and CI/PR gating via a GitHub Action, with optional hosted run history and diffs while keeping local evaluation artifacts private by default.
- Pieces
- CLI · capture SDK · GitHub Action · hosted app
- Providers
- Anything LiteLLM supports
- Quality
- ~2,200 tests · 90% coverage floor in CI
- Privacy
- Local-first — nothing uploads until you push
- License
- Apache-2.0 (CLI) · MIT (SDK)
Recently shipped
- 1.1.0Sep 19, 2026The migration policy travels inside every pushed run, so the hosted PR gate checks the same budgets as the local verdict.
- 1.0.0Sep 18, 2026First stable release, with a written SemVer contract.
allbecamecompare. - 0.15.0Sep 15, 2026Relicensed to Apache-2.0; the capture SDK is MIT.
- 0.14.0Sep 10, 2026Teacher-forced multi-round replay and repeated sampling per example.
uv pip install evalshift
evalshift init --ci
evalshift compare --suite-name support_agent --to <candidate>
Cartridge B — KidViq
Independent product · built and shipped · Jan 2024 – Nov 2025
Daycare management, as easy as playing.
An all-in-one platform for daycares and kindergartens: admins run the whole center from the web, teachers and parents stay in sync from their phones. Used in production across hundreds of kindergartens.
What it covers
Three apps, one platform
- Admin
- Web app — React · TypeScript
- Teachers
- iOS & Android — Flutter · Dart
- Parents
- iOS & Android — Flutter · Dart
- Backend
- Node.js · PostgreSQL · AWS
What I built
- Built and shipped a daycare/kindergarten management platform used in production across hundreds of kindergartens, covering scheduling, payments, child-records CRM, documents, messaging, and development tracking.
- Led a small team across a React/TypeScript web platform and two Flutter apps for parents and teachers; owned architecture, code-quality standards, and delivery across Node.js, PostgreSQL, AWS, and Flutter/Dart.
Specifications
What's inside.
Three layers, one engineer. Hover a part to find it in the exploded view.
Face plate & controls
Product Engineering
The part people touch: fast, intuitive, maintainable product surfaces.
- TypeScript
- JavaScript
- React
- Next.js
- Node.js
- REST APIs
- WebSockets
- Component systems
At a glance
- Experience
- 8+ years
- Role
- Senior product engineer · team lead
- Based in
- St. Petersburg, FL · US Eastern
- Degree
- B.Sc. Mathematics & Computer Science
Logic board
Applied AI & Data
LLM systems you can measure — and trust.
- LLM evals
- Agent/tool-call evaluation
- LLM-as-judge
- RAG
- Embeddings
- pgvector/vector search
- PostgreSQL
- SQL
Power unit
Cloud & Delivery
Shipping it, safely, end to end.
- AWS ECS/EKS
- Lambda
- API Gateway
- S3
- IAM
- DynamoDB
- CDK
- Docker
- CI/CD
- Python
- Go
- FastAPI
Contact
Let's build the next one.
Email is the fastest way to reach me. I’m also on LinkedIn and GitHub.