What langfuse/langfuse shipped
Written by FoxPlug from public releases; not affiliated with Langfuse. An automatic summary of the public release, pull request and commit data of github.com/langfuse/langfuse. Langfuse did not write it and does not use or endorse FoxPlug. Every line links to the public change it describes.
Get a weekly update like this for your product, free
Week of September 21, 2026
What shipped
- v4.46.0 released with experiment comparison improvements, dataset formatting, skill management, and larger page size options. Release
- Experiment side-by-side comparison now defaults to one column per experiment with collapsible summary and shared score details. Pull request #17929
- Basic skill management endpoints added to the REST API for creating, reading, and updating skills. Pull request #17803
- Dataset items now render formatted JSON in code editors for input, expected output, and metadata fields. Pull request #17932
- Dataset run-items reads routed to read replica for improved performance on v3 API. Pull request #17919
- Dashboard query parameters with large filter values now sent as multipart body to avoid proxy rejections. Pull request #17970
- Decision-model evaluators now correctly link to their scores, fixing empty scores tables. Pull request #17954
- Trace playback feature removed from trace navigation header. Pull request #17927
- Topics traces now render and measure dedicated trace transcripts with fixed sections for facts, tools, and state. Pull request #17925
- Scores list query performance improved by deduplicating events rows via GROUP BY and INNER JOIN. Pull request #17883
Why it matters
This week brings significant UX refinements to experiment comparison and dataset viewing, new skill management capabilities, and multiple performance optimizations for dashboard queries and data retrieval. The changes make core workflows faster and cleaner, with better defaults and improved data presentation throughout the interface.
Changelog entry
- feat(experiments): streamline side-by-side comparison by defaulting to one column per experiment with collapsible summary Release
- feat(web): add formatted json display to dataset items for input, output and metadata Release
- feat(skills): basic skill management via REST endpoints for CRUD operations Release
- feat(web): add 100-row page size option to tracing tables Release
- fix(dashboard): send large query params as multipart body to avoid proxy rejections Pull request #17970
- fix(evals): link decision-model evaluators to their scores by evaluator ID Pull request #17954
- perf(datasets): route v3 dataset-run-items reads to the read replica Pull request #17919
- perf(scores): dedup events rows via GROUP BY and INNER JOIN for faster list queries Pull request #17883
- fix(web): remove trace playback feature and related controls Pull request #17927
- feat(topics): render and measure dedicated trace transcripts with fixed sections Pull request #17925
- fix(experiments): hide cost column by default in experiment comparison grid Pull request #17947
- fix(web): reduce tooltip open delay from 700ms to 300ms app-wide Pull request #17946
- fix(web): update breadcrumb styling with thinner slashes and smaller chevrons Pull request #17934
v4.46.0 ships cleaner experiment comparisons, formatted dataset JSON views, skill management APIs, and performance improvements across dashboard queries and data retrieval.
v4.46.0 is out. Experiment side-by-side comparisons now feature collapsible summaries and better score details. Dataset items display formatted JSON for easier viewing. New REST endpoints enable skill management. Performance optimizations route dataset queries to read replicas and fix dashboard query parameter handling. Plus UI improvements throughout including faster tooltips and better breadcrumb styling.
Week of September 14, 2026
What shipped
- Thread API now returns conversation history and current turn per thread for better transcript organization. Pull request #17613
- Gateway generation metadata reorganized into consistent namespaces with proper HTTP status prefixes. Pull request #17645
- Every ingestion event now authorized through the policy core for improved access control. Pull request #16715
- Automations can now filter prompt events by label with live label loading in the UI. Pull request #17612
- Trace-level scores moved from tree visualization to the trace header for better visibility. Pull request #17642
- Table rows now clear when switching between saved views, filters, or searches to prevent stale data. Pull request #17692
- JSON viewer restyled with unified table appearance across the application. Pull request #17402
- Trace batch dispatcher now measures pending trace size estimates and reports batch size distributions. Pull request #17586
- Capacity experiment telemetry added for trace batch scaling to track admitted volume and queue snapshots. Pull request #17614
- Web processes publish event-loop delay metrics for CloudWatch autoscaling. Pull request #17632
Why it matters
This week includes structural improvements to core data APIs like threads and gateways, better trace observability with metrics and UI reorganization, and expanded automation capabilities. Infrastructure changes improve event-loop monitoring and batch processing visibility for production deployments.
Changelog entry
- feat: conversation history and current turn now returned per thread Pull request #17613
- fix(ai-gateway): reorganized generation metadata into consistent namespaces Pull request #17645
- feat(auth): all ingestion events authorized through policy core Pull request #16715
- feat(automations): label filtering added for prompt events in GitHub dispatch and webhooks Pull request #17612
- fix(web): trace-level scores moved from tree visualization to trace header Pull request #17642
- fix(table): stale rows cleared when switching filter scope Pull request #17692
- fix(web): JSON viewer restyled with unified table appearance Pull request #17402
- feat(web): event-loop delay metrics published for CloudWatch autoscaling Pull request #17632
- feat(trace-batching): pending trace size estimates measured and reported Pull request #17586
- fix(logging): severity field added to JSON logs for GCP Cloud Logging compatibility Pull request #17606
This week: thread transcripts with conversation history, gateway metadata reorganization, label-based prompt automation filters, and improved trace-level visibility. Full details at the preview links.
This week's releases focus on better data organization and observability. Threads now return full conversation history per turn, gateway generation metadata is consistently namespaced, and automations support label-based filtering. Infrastructure improvements include event-loop delay metrics for autoscaling and batch processing visibility. Check the preview links for details.