Data42Data42

Services

Everything between raw data and a working AI product.

We can own a whole platform end to end, or plug senior engineers into the gaps in your team. Most clients start with one of the six areas below and expand from there.

Build the foundation

01

Data Engineering

Pipelines, warehouses, and lakehouses on Snowflake, BigQuery, Databricks, and AWS. Airflow and dbt with tests, CI, and lineage, so a broken source fails loudly at 3am instead of quietly in next quarter's board deck. We also do the unglamorous half: migrating off legacy ETL and cutting warehouse spend.

  • Batch and streaming pipelines with Airflow, Spark, Kafka, and AWS Lambda, Step Functions, Glue, and EMR
  • Warehouse and lakehouse design on Snowflake, PostgreSQL, Databricks, and S3 with Iceberg tables
  • dbt and Snowpark for transformations, with tests, documentation, lineage, and CI on every pull request
  • Migrations we have run more than once: Hadoop to AWS, legacy ETL and SSIS to Matillion and Snowflake
  • Database development where it still matters: Oracle PL/SQL, PL/pgSQL, T-SQL, query optimisation
  • Cost work: Glue to Lambda where the workload is small, warehouse sizing, partitioning, file formats

02

Analytics Engineering & BI

One definition of revenue, churn, and active customer, agreed once and enforced in code. Tested, documented dbt models, a semantic layer your BI tool reads from, and dashboards in Apache Superset, Power BI, Metabase, or Looker that analysts can change without filing a ticket.

  • Metric definitions and a dbt semantic layer shared across the company
  • Tested, documented dbt models with clear ownership
  • Apache Superset as a product: embedded dashboards, row-level security for multi-tenant data, Jinja-templated charts
  • Power BI, Metabase, and Looker when they fit your licensing and your analysts better
  • Brand, market, and influencer analytics for research and marketing teams

Build on top of it

03

AI Engineering

Retrieval over your documents and databases, agents that do real work, and LLM features inside your product. We build the evaluation harness before the feature, so you can tell whether a change made the model better or only different. Claude, OpenAI, and open models, with cost and latency budgets from day one.

  • Hybrid retrieval over documents, warehouses, and internal systems: vector search plus keyword matching, tuned against a test set rather than a demo
  • Ingestion for AI: social media, PDFs, podcasts and video through transcription, news, and academic literature, normalised into one searchable store
  • Agents and workflow automation on Claude, OpenAI, and open models, orchestrated with Step Functions, with circuit breakers so a failing step cannot run forever
  • MCP servers that give Claude read-only, scoped access to your databases and AWS accounts, with credentials redacted before anything reaches the model
  • Eval harnesses, guardrails, and cost control: real examples, graders, token budgets, multi-provider fallback, cheaper models where they are good enough

How an engagement runs

  1. 01

    Scope

    Three days with your team and your data. You get the two or three use cases worth building, what each needs from your warehouse, and an estimated monthly running cost. If none clear the bar, we say so.

  2. 02

    Evals first

    Weeks one and two. Before the feature, the measurement: a labelled example set and graders built with your domain experts, and a published baseline number.

  3. 03

    Build

    Weeks three to eight, in your cloud and your repo. A weekly demo against the eval set, so better is a number rather than an opinion, with accuracy, latency, and cost per request side by side.

  4. 04

    Handover

    Runbook, dashboards, alerting, and a working session with your engineers. The evals ship with the code and keep running after we leave.

AI-assisted engineering, set up properly

Your developers already have Claude Code licences. The gap is everything around them. A faster first draft does not fix a slow review process; it loads more onto the people doing the reviewing. So we run adoption as a six-week programme inside your repositories, not a workshop: conventions files per repo, permission and secrets policy your security team signs off, usage metrics with a baseline before anything changes, real tickets paired with your engineers, and a written review workflow for diffs nobody typed. We run our own work this way, including this website.

04

Data Products & Software

The application around the data: internal tools, customer-facing portals, APIs. TypeScript, React, Next.js, and Python on AWS, with infrastructure as code. Built by the same people who built the pipeline underneath, so nobody argues about whose bug it is.

  • Web applications and APIs in Python and Django, TypeScript, React, and Next.js
  • Backends on AWS with CDK, containers, and no SSH: SSM-based access and deployments, one-command releases
  • Data products: embedded dashboards, client portals with per-tenant data segregation, insight tools
  • Security by layers: route guards, backend middleware, and database-level checks, with secrets scrubbed from logs

Add people to your team

05

Embedded Engineers

Named senior data and AI engineers who join your team, your repo, and your standup, and stay long enough to know your data model better than your last hire did. No juniors learning on your budget, and no CVs you have to interview twice.

  • Senior engineers embedded in your team, reporting to your leads, in your Jira, your Git, and your Slack
  • Accounts and repo access on day one; first merged pull request inside two weeks
  • Working hours that cover a full European day, with a live window for US East Coast teams
  • NDA and IP assignment before we see a line of your code; used to access reviews and security questionnaires
  • Fully outsourced delivery when you would rather hand over a platform or a migration end to end

Three shapes

01

Embedded engineers

One to six engineers work only on your roadmap, under your tech lead, in your process. You direct the work day to day. Billed monthly per engineer.

02

Managed delivery

You describe the outcome and we own the plan, the team, and the date. One named lead our side, one decision-maker yours.

03

Senior retainer

A fixed block of senior hours each month for architecture review, unblocking, and the work with no owner. For teams that need judgement more than hands.

Why Armenia

Armenia runs on UTC+4 all year, so a working day in Yerevan overlaps a European one without anyone taking a call at six in the morning, and most EU capitals are a four-hour flight away. The market is small enough that we personally know most of the senior data engineers worth hiring, which is the real reason our people stay on the same account for years. We are not the cheapest place to buy an engineer and we do not compete there. You are paying for one senior person who stays, not a discount on a junior who rotates out in eight months.

06

Data Engineering Academy

A twelve-week path that takes an analyst who writes good SQL and turns them into someone who can own a pipeline in production. Taught on real projects instead of toy datasets, reviewed line by line by a senior engineer. Run it for your own team on your own stack, or hire the engineers who finish it.

  • For analysts blocked by pipelines, backend engineers moving into data, and juniors with nobody senior to review them
  • Not for complete beginners: you should be comfortable in SQL and able to read Python
  • Cohorts capped at fifteen people so code review is real
  • Outcome: a working repository you keep, not a certificate

The path

  1. 01

    SQL that survives production

    Window functions, incremental logic, query plans, and why the query that runs on your laptop times out on the warehouse.

  2. 02

    Python for data engineers

    Packaging, typing, testing, environments. Code someone else can run in six months.

  3. 03

    Modelling

    Dimensional and wide-table designs, slowly changing dimensions, grain, and choosing what the business actually asked for.

  4. 04

    dbt

    Projects, sources, tests, docs, macros, and CI on every pull request.

  5. 05

    Orchestration with Airflow

    DAG design, idempotency, backfills, retries, and what to do at 3am when a run fails.

  6. 06

    The warehouse and the bill

    Snowflake, BigQuery, and Databricks internals that move your invoice: partitioning, clustering, sizing, file formats.

  7. 07

    Ingestion and streaming

    APIs, change data capture, files, and a first look at Kafka. When streaming is the answer and when it is an expensive way to do batch.

  8. 08

    Quality, observability, and cost

    Tests versus monitors, freshness and volume checks, lineage, alert fatigue, and reading a cloud bill.

  9. 09

    Capstone

    A working pipeline from source to dashboard, reviewed line by line, in a repo the participant keeps. Optional module on using Claude Code on a data codebase.

Formats

01

Cohort

Twelve weeks part-time, about eight hours a week: two live sessions plus exercises. Capped at fifteen people.

02

Corporate

The same path rebuilt around your stack, your warehouse, and your naming conventions. Four to twelve weeks, remote or on-site.

03

Mentoring

One senior engineer, one or two of your juniors, weekly: pairing, code review on real pull requests, and a plan for the next quarter.

The academy is also how we hire

Strong graduates get offered a place with us. Clients who train a team with us can hire the engineer who taught them, or take one of our people on while their own bench catches up. If you are building a data team in Armenia, we will train your juniors and tell you honestly which ones to keep.

How an engagement runs

Four steps, each with a real duration.

  1. 01 · 30 minutes

    Technical call

    You describe the stack, the deadline, and what has already been tried. You talk to the engineer who would do the work, not a salesperson. If we are not the right fit, we say so and point you elsewhere.

  2. 02 · 1 to 2 weeks

    Discovery

    We read the code, the models, and the tickets. You get a written assessment and a scoped plan with an estimate you could hand to another vendor.

  3. 03 · Weekly demos

    Build

    Work ships in small reviewable pull requests inside your repo. Tests, documentation, and infrastructure as code are part of every deliverable, not a phase at the end.

  4. 04 · Your call

    Handover or stay

    Once your team can own what we built, we step back. Or an engineer stays embedded for as long as the work needs one. Either way, everything we built is yours.

Fit

We are not the right choice if

  • You want a slide deck about AI strategy and no code.
  • You need twenty engineers by next month.
  • You are looking for the cheapest hourly rate.

If none of those describe you, the next step is a call with the engineer who would do the work.