What langwatch/langwatch shipped
Written by FoxPlug from public releases; not affiliated with Langwatch. An automatic summary of the public release, pull request and commit data of github.com/langwatch/langwatch. Langwatch did not write it and does not use or endorse FoxPlug. Every line links to the public change it describes.
Get a weekly update like this for your product, free
Week of September 21, 2026
What shipped
- LangWatch 3.18.0 ships with a Helm chart for self-hosted deployments. Release
- The CLI now runs from a committed launcher instead of a bundled hook, and agent testing can compare agents in one run. Release
- A new integration guide for LiteLLM Proxy lets teams trace every call through their internal LLM gateway without touching application code. Pull request #8309
- Datasets can now hold uploaded images and files in cells, and those attachments reach models and agents mapped to those columns. Pull request #7964
- Connected self-hosted installs can now call LangWatch-hosted services with usage metered and capped against the license. Pull request #8232
- Fixes for upgrading a self-hosted 3.16 Helm install to 3.18 on EKS were validated on a real customer deployment. Pull request #8323
- A new docs page guides administrators through moving to their own identity provider with self-serve updates. Pull request #8271
- When Instant Evals are disabled, the UI now prompts to enable them instead of showing a model popover. Pull request #8295
- A search now correctly judges sentences classified as judgements, and Test Connection sends real calls. Pull request #8268
- Every tenant-scoped Postgres model is now automatically derived into the LangWatchQL catalog, matching the ClickHouse approach. Pull request #8209
Why it matters
3.18.0 brings self-hosted deployments to Kubernetes with a Helm chart, adds support for images and files in datasets, and enables self-hosted instances to use LangWatch-hosted services with metered usage. Teams using LiteLLM as their gateway can now trace all calls without code changes. These updates address deployment, extensibility, and integration gaps.
Changelog entry
- Helm chart for LangWatch application deployed to Kubernetes [0] Release
- CLI runs from a committed launcher instead of bundled hook [1] Release
- Agent testing can compare agents in one run [1] Release
- LiteLLM Proxy integration guide for tracing without application code changes [8] Pull request #8309
- Datasets support uploading and storing images and files in cells [21] Pull request #7964
- Uploaded attachments are delivered to prompts and agents mapped to columns [21] Pull request #7964
- Connected self-hosted deploys can call LangWatch-hosted services with metered usage [30] Pull request #8232
- Self-serve update guide for moving to your own identity provider [18] Pull request #8271
- Instant Evals prompt displays when flag is off instead of model popover [11] Pull request #8295
- Fixes validated for upgrading self-hosted 3.16 Helm installs to 3.18 on EKS [4] Pull request #8323
LangWatch 3.18.0: Helm chart for self-hosted K8s, images and files in datasets, LiteLLM Proxy integration, connected self-hosted with metered hosted services.
LangWatch 3.18.0 is out. New Helm chart for self-hosted Kubernetes deployments. Datasets now support images and files. LiteLLM Proxy integration lets teams trace calls through their internal gateway without code changes. Connected self-hosted installs can use LangWatch-hosted services with metered usage against the license.
Week of September 14, 2026
What shipped
- Instant Eval runs execute a LangWatchQL statement as a judgment job with progress tracking and persisted results, enabling bulk evaluation workflows beyond the exploratory inline query cap. Pull request #8208
- The Instant Eval CLI waits for answers, accepts target shorthand syntax, and estimates spend before running, making evaluation accessible without writing LangWatchQL by hand. Pull request #8216
- Instant Eval judgments meter on the gateway spend spine as Stripe charges with a 1 USD free budget, moving billing onto the standard cost tracking system. Pull request #8220
- LangWatchQL gains eval functions that call the classifier interface, letting judgment calls sit inline in analytics queries alongside metrics and dimensions. Pull request #8201
- Any LangWatch API key now has one query door into LangWatchQL via POST /api/v1/query, running read-only SELECT statements across all projects the key can read with ClickHouse row policies enforcing scope. Pull request #8113
- Custom dashboard charts now run in their own sandboxed frame on a separate route, allowing widgets to import any module under the production Content-Security-Policy. Pull request #8151
- Sign-in now offers one-click social authentication through Google, GitHub, and Microsoft buttons alongside Auth0, removing the double-SSO-click from production. Pull request #8143
- Trace reading logic (conversation threading, LLM span splitting, digest generation) moved from React components into reusable extraction modules for use across the platform. Pull request #8194
- Ingestion pull outcome writes now carry typed dispatches instead of casts, improving bookkeeping accuracy for customer assignment and timestamps. Pull request #8111
- Cost rollup checks moved from a nightly cron to an event-driven per-tenant process that triggers only when cost changes land, replacing batch checking with reactive updates. Pull request #8121
Why it matters
Instant Evals shipped as a complete feature this week: runs, CLI, judgment functions, billing, and metrics. LangWatchQL now has a unified query door that any API key can use. These changes let users evaluate bulk text through the platform and access analytics programmatically without building custom integration.
Changelog entry
- feat(instant-evals): the Instant Eval run, a judgment job over an LWQL statement with progress and persisted judgments Pull request #8208
- feat(instant-evals): the CLI that waits for the answer, the target shorthand and estimate before spend Pull request #8216
- feat(instant-evals): meter judgements on the gateway spend spine, a Stripe meter and a 1 USD free budget Pull request #8220
- feat(lwql): eval functions judged by the classifier interface (Instant Evals) Pull request #8201
- feat(query): self-describing LangWatchQL door + whoami --json Pull request #8113
- feat(analytics): serve chart sandbox frame on its own route so widgets can import any module Pull request #8151
- feat(identity): one-click social sign-in — the Auth0 connection bridge, native-provider groundwork, and AUTH_PROVIDER Pull request #8143
- refactor(traces): framework-free conversation, chat split, LLM span and bounded digest modules Pull request #8194
- feat(governance): replace the cost-rollup comparator cron with an event-driven per-tenant costRollupWatch process manager Pull request #8121
- fix(instant-evals): review sweep over the ten Instant Evals PRs Pull request #8230
- fix(instant-evals): end-to-end dogfood on main, five fixes in how numbers and words reach the caller Pull request #8233
- fix(clickhouse): renumber the coding-agent usage migration to 00099, 00097 was taken Pull request #8235
Instant Evals are live: run bulk judgments through LangWatchQL, use the CLI without writing queries, and only pay for what you evaluate. One query door for any API key. Dashboard charts now run in sandboxed frames.
Instant Evals shipped this week with three layers: LangWatchQL now has eval() functions for inline judgments, runs execute bulk evaluation jobs with progress tracking, and the CLI makes it accessible without SQL. Billing integrates with Stripe on the standard cost spine. Separately, any LangWatch API key now has one query door into LangWatchQL—POST /api/v1/query runs read-only SELECT across all readable projects. Custom dashboard charts moved to sandboxed frames. Sign-in gained one-click social auth.