1. 03 Sep, 2026 4 commits
    • feat(plugin): add MiniMax-H3 /v2 video generation to the hailuo task … (#7168) · aece11d2
      * feat(plugin): add MiniMax-H3 /v2 video generation to the hailuo task plugin
      
      MiniMax-H3 speaks a different contract from the other Hailuo models, so the
      hailuo task plugin now branches on the upstream model instead of adding a Go
      adaptor:
      
      - submit builds /v2/video_generation with a multimodal `content` array
        (text, first/last frame images, reference video/audio, or a full
        `metadata.content` passthrough), an explicit `ratio`, and 768P/2K
        resolutions; `metadata.callback_url` and `metadata.aigc_watermark` pass
        through
      - query uses /v2/query/video_generation/{task_id} and parses the
        `{"task": {...}}` envelope, falling back to the /v1 shapes for every other
        model
      - the /v2 result is a public CDN URL, so its artifact is proxied
        credentialless instead of through /v1/files/download
      - request bounds (duration 4-15, resolution 768P/2K, ratio whitelist, at most
        2 frame images and 9/3/3 reference images/videos/audios) are enforced while
        the request body is built, which the host runs during validation, so an
        out-of-range duration is rejected with a 400 before it can become a billing
        multiplier
      - duration and resolution are reported as usage facts only. Like the rest of
        this plugin, extractUsage returns no billing ratios, so per-call pricing is
        flat and 2K/duration pricing is expressed through the model's tiered billing
        expression over those facts.
      
      Query hooks are driver hooks and are documented to receive `ctx.model` and
      `ctx.upstreamModel`, but polling has no relay info and never populated them.
      The polling and realtime-fetch call sites now carry the persisted task model
      properties and the plugin adaptor maps them onto the query context, with
      `upstreamModel` falling back to the origin name for tasks submitted without a
      channel mapping.
      
      * fix(plugin): validate Hailuo H3 requests and errors
      星云猫 committed
    • fix(relay): follow-up billing integrity and conversion completions (#7170) · bbd97446
      Deferred follow-ups from the relaykit-tools review cycle, verified by
      live end-to-end billing tests:
      
      - billing: normalize Gemini modality keys consistently between stream
        merge and settlement (case/whitespace variants no longer drop
        independent audio/image pricing) and sum duplicate modality entries
        on both paths
      - billing: sync legacy flat Claude cache-creation fields from the
        CacheCreation sub-object (including zeroing) and fall back to flat
        fields only when the snapshot never carried a sub-object, closing a
        stale 1h-cache overcharge path in cascaded deployments
      - relay: move Chat-to-Claude and Chat-to-Gemini stream conversion state
        from gin.Context onto RelayInfo and reset it with SendResponseCount in
        InitChannelMeta, so channel retries start clean while per-request
        state (stream error collection, conversion diagnostics, channel
        chain, billing accumulators) survives
      - relay: Claude channel now serves Gemini-format clients (request via
        registry conversion, response and stream composed through the Chat
        pivot), removing the last unimplemented conversion direction
      - relaykit: recognize legacy pseudo tool names (googleSearch,
        codeExecution, urlContext) in the toolconv decode stage and drop the
        string-matching bypass in the Chat-to-Gemini converter; native Gemini
        tool output is restored and non-Gemini targets follow standard loss
        diagnostics
      - relaykit: attach upstream Gemini usage (with billing_usage sidecar)
        to intermediate stream chunks so converted Claude streams report
        upstream truth from message_start, and preserve the sidecar through
        Claude stream usage merges; billing settlement unchanged
      - billing: clamp negative Total-Prompt completion derivation, OR the
        Estimated flag across cross-dialect snapshot replacement, and fill
        canonical OpenAI prompt details via field-wise merge
      Calcium-Ion committed
  2. 01 Sep, 2026 2 commits
    • feat(relay): hosted-tool conversion fidelity, reasoning normalization, and… · 0ed497f0
      feat(relay): hosted-tool conversion fidelity, reasoning normalization, and billing usage integrity (#7137)
      
      * feat(relaykit): preserve hosted tools across conversions
      
      - add protocol-neutral hosted-tool DTOs, conversion metadata, and loss policies
      - bridge citations, grounding metadata, and hosted-tool stream lifecycles
      - document the public conversion behavior and channel policy controls
      
      * refactor(relaykit): normalize reasoning and thinking intent
      
      - centralize provider-neutral reasoning intent, effort, and budget mappings
      - parse model suffixes at the host entry boundary while preserving provider-owned tails
      - keep adaptive Claude thinking and explicit zero-token compatibility consistent
      
      * fix(billing): preserve authoritative usage across relay hops
      
      - carry native BillingUsage sidecars through direct and streamed protocol bridges
      - merge partial and terminal usage monotonically with safe fallback settlement
      - retain cache metadata, penultimate usage, and per-call Gemini tool surcharges
      
      * feat(relay): bridge Responses with Claude and Gemini protocols
      
      - add direct request, response, and stream converters across supported relay formats
      - expose Claude count_tokens and Chat-to-Responses compatibility endpoints
      - carry conversion diagnostics through the host while retaining the curated public goldens
      
      * fix(relay): wire relaykit conversions into host channels
      
      - connect handlers, adaptors, and channel settings to the standalone conversion layer
      - keep model mapping, pricing identity, retries, and provider-specific suffix behavior aligned
      - ignore local audit artifacts and retain focused public regression coverage
      Calcium-Ion committed
  3. 31 Aug, 2026 3 commits
  4. 30 Aug, 2026 12 commits
    • fix initialize database · 74158715
      CaIon committed
    • fix(relay): bound the wait for upstream response headers (fixes unbounded heap… · b518d003
      fix(relay): bound the wait for upstream response headers (fixes unbounded heap growth → OOM) (#6949)
      
      * fix(relay): bound the wait for upstream response headers (fixes unbounded heap growth)
      
      The relay transport sets a dial timeout, a TLS handshake timeout and an expect-continue
      timeout, but nothing bounds how long it waits for the upstream *response headers* after
      the request has been written. An upstream that accepts the connection and then never
      answers -- without sending FIN/RST, which is what happens when a NAT/firewall silently
      drops the flow or the provider hangs -- parks the goroutine in
      net/http.(*persistConn).roundTrip forever.
      
      That goroutine keeps the whole request alive, which in practice means three copies of the
      request body stay reachable for the lifetime of the process: the raw bytes from
      io.ReadAll in CreateBodyStorageFromReader, the decoded messages held as json.RawMessage,
      and the re-marshalled upstream body from common.Marshal. BodyStorageCleanup cannot help
      here: it runs after c.Next() returns, and for these requests c.Next() never returns.
      
      Measured on v1.0.0-rc.23 in production (see #6947 for the full evidence):
      
        - 23 goroutines stuck in persistConn.roundTrip on a single 40h-old instance,
          blocked between 353 and 1894 minutes (5.9h to 31.5h)
        - 96.9% of the live heap, sampled after a forced GC, attributable to those three
          body copies (HeapAlloc 892 MiB surviving three GC cycles; HeapObjects dropping
          30x while bytes dropped only 25%)
        - the live floor grows with uptime: 33.7 MiB at 0.1h, 89.2 at 13.8h, 510.0 at 40.1h,
          955.2 at 146.8h, OOMKilled at 172.9h -- same image, same config, same load
      
      Doubling the memory limit and adding GOMEMLIMIT only moved the OOM from 132h to 172.9h.
      
      RELAY_TIMEOUT (http.Client.Timeout) cannot be used for this: it covers the whole response
      read and would cut legitimate long streaming calls, which is why it defaults to 0.
      ResponseHeaderTimeout only bounds the wait for the headers; streaming after they arrive is
      unaffected.
      
      The default is deliberately generous. Non-streaming upstreams usually send the response
      headers only once generation has finished, so the value has to leave room for a long
      completion. 1800s is 12x shorter than the shortest hang observed here while leaving
      several times the headroom a normal non-streaming request needs; 0 restores the previous
      unbounded behaviour.
      
      The assignment goes next to the other transport.* lines rather than inside the else
      branch: newRelayHTTPTransport() normally takes the http.DefaultTransport.Clone() path,
      and DefaultTransport does not set ResponseHeaderTimeout either.
      
      This repo already sets ResponseHeaderTimeout on its other outbound transports
      (controller/model_sync.go, controller/ratio_sync.go); the relay path appears to have
      been missed.
      
      Refs #6947. Likely also the root cause of #6731, which reported the same symptom
      (production OOM on /v1/responses after ~64h) but was closed for template reasons.
      
      * review: clamp overflowing timeout values and switch the test to testify
      
      Addresses the two CodeRabbit findings on this PR.
      
      Overflow (common/init.go:113): a RELAY_RESPONSE_HEADER_TIMEOUT beyond ~9.2e9 seconds
      overflows time.Duration and can wrap into a *tiny positive* timeout, which would cut
      every relay request instead of only the stuck ones. The value is now clamped before the
      conversion, with regression tests for both the negative and the overflowing input.
      
      I did not add fail-on-startup validation for negative values, for two reasons: the
      existing `if seconds > 0` guard already treats them as "disabled", and the neighbouring
      env-driven timeouts in this file are less strict still -- RelayIdleConnTimeout is
      converted with no guard at all. Failing startup on a bad value would be a behaviour
      change out of step with the rest of the file; happy to add it if you'd prefer that
      direction repo-wide.
      
      Test style: switched to testify (require.Equal / require.Zero / require.Positive), which
      is what every other test under service/ uses.
      
      go build, go vet and go test ./common/... ./service/... pass.
      (`go build ./...` fails on the `web/dist` embed both with and without this change -- the
      frontend bundle is not checked in.)
      txgo committed
    • fix(model): return string from JSON column Valuers for pg simple protocol · 6eb6f35e
      With PrepareStmt disabled, PostgreSQL queries run over pgx's simple
      protocol, which encodes every []byte parameter as a bytea hex literal
      ('\x...'). driver.Valuer implementations returning []byte from
      json.Marshal therefore fail json-column writes with SQLSTATE 22P02
      (reported on the channels UPDATE path via ChannelInfo).
      
      Reproduced against a live PostgreSQL 16: []byte Valuer into a json
      column fails under simple protocol, string succeeds; []byte into a
      text column silently stores the hex literal (no such path exists in
      the repo today — audited all Valuers, json.RawMessage fields, and raw
      SQL call sites).
      
      - ChannelInfo, Properties, TaskPrivateData, JSONValue Value() now
        return string; zero-value nil semantics unchanged. Task.Data
        (bare json.RawMessage) is unaffected — database/sql's default
        converter already passes it as expected.
      - Their Scan() counterparts now accept both []byte and string via a
        shared jsonScanBytes helper: SQLite returns string for these columns
        once Value() emits string, and the old []byte-only assertions
        silently zeroed the field (caught by the model test suite).
      - Add regression tests locking both contracts: json-column Valuers
        must return string (or nil for zero values), Scanners must accept
        []byte and string.
      
      Verified end-to-end against PostgreSQL 16 with the real model types:
      Channel create/update/read-back, Task json fields, PrefillGroup items.
      CaIon committed
    • fix(sqlite): enable WAL + working busy timeout + _txlock=immediate to stop… · 1751f43e
      fix(sqlite): enable WAL + working busy timeout + _txlock=immediate to stop concurrent write lockouts (#7030)
      
      * fix(sqlite): enable WAL + working busy timeout + _txlock=immediate to stop concurrent write lockouts
      Xayinn committed
    • fix(subscription): 无有效订阅时前端如实显示「仅用订阅」偏好 (#6222) (#7086) · b5b94bc6
      Co-authored-by: Claude <noreply@anthropic.com>
      ruiyunzhao committed
    • feat(web): factory task plugins update only with the system · dc4732cf
      Marketplace install/upgrade on a factory-served plugin actually created
      a permanent override shadowing every future built-in release. The card
      now shows an informational "Updates with the system" badge instead of
      the action, while keeping the built-in vs marketplace version line and
      the upgradable state badge visible. Deliberate overrides are untouched:
      upload and marketplace actions on overridden or third-party plugins
      behave as before, and the plugins table now hints when an override
      lags behind the shipped built-in version so operators know deleting it
      restores the newer factory plugin.
      CaIon committed
    • fix(model): disable PostgreSQL prepared statements for pooler compatibility · 66031a09
      GORM v1.25.2 closes cached prepared statements asynchronously on any SQL
      error and immediately re-Parses the same deterministic name (pgx's
      stmt_<sha256>) on the same client connection. Transaction-pooling proxies
      (PgBouncer >=1.21 with max_prepared_statements, Neon, Supabase) respond
      with FATAL "prepared statement name is already in use" (SQLSTATE 08P01)
      and drop the connection. PreferSimpleProtocol only disables pgx's
      implicit prepare and never covered GORM's explicit PrepareStmt cache.
      
      - PostgreSQL now runs with PrepareStmt disabled entirely; named prepared
        statements are fundamentally session state and cannot be made safe
        under transaction pooling. Parse/plan cost is noise for this workload.
      - Upgrade gorm to v1.25.12 so MySQL/SQLite statement caches (still
        enabled) no longer churn close/re-prepare on ordinary SQL errors;
        v1.25.9+ restricts eviction to driver.ErrBadConn. Deliberately not
        v1.26+, whose LRU eviction has an open use-after-close race (#7831).
      - sanitizeDBError now attaches a remediation hint on 08P01/42P05 so
        affected deployments can self-diagnose from the log line.
      CaIon committed
    • feat(task): resolve channel-mapped aliases and case variants for plugin models · 6c22550e
      Channel model_mapping keys exposed in a channel's model list now act as
      first-class aliases for task-plugin models across the whole line:
      
      - Derived alias view (model/task_model_alias.go): built from enabled
        channels' model_mapping, chain-following with cycle detection, declared
        names always win, cross-plugin conflicts dropped. Rebuilt on channel
        cache refresh, registry generation change, and a 60s TTL.
      - Request path: PinTaskPluginEndpoint resolves declared-name case folds
        and mapping aliases before endpoint lookup (never rewriting the body
        until the endpoint is claimed), pins with MappedModel, and the decode
        contract accepts alias echoes without loosening model ownership for
        normal pins. Legacy /v1/tasks submit folds case variants the same way.
        Fixes aliases on POST /v1/responses silently falling through to the
        main relay against task channels.
      - Mapping order: ModelMappedHelper now runs before the plugin submit
        hook builds and caches the upstream body, so channel model_mapping
        actually reaches the upstream request. Plugins receive the mapped
        name as ctx.upstreamModel in both decode and submit contexts.
      - Billing: identity stays the origin name; when the alias has no tiered
        expression, the selected channel's mapping tail expression applies.
        Pricing page and billing-expr smoke tests resolve aliases to the
        owning plugin's usage schema.
      - Case folding: ASCII-only fold with exact-match priority; same-plugin
        and cross-plugin fold collisions rejected at registration.
      - Plugins: model-keyed rate tables, req_key derivation, and combo
        validation in doubao/kling/jimeng/hailuo/vidu/sunoapi now key on
        ctx.upstreamModel || ctx.model; render/echo paths keep ctx.model.
      CaIon committed
  5. 29 Aug, 2026 9 commits
  6. 27 Aug, 2026 3 commits
  7. 26 Aug, 2026 3 commits
  8. 21 Aug, 2026 1 commit
  9. 18 Aug, 2026 3 commits
    • feat(web): fade in streamed response words and harden playground editor (#6895) · 137d1171
      * feat(web): fade in newly streamed response words
      
      Animate only new word-level deltas while markdown is still streaming, and
      cache markdown-it instances per parser id so concurrent Response trees do
      not rebuild or reparse on every render.
      
      * fix(web): keep CodeMirror editor alive across keystroke re-renders
      
      Deliver onKeyDown through a ref instead of the extensions memo so a new
      handler identity no longer tears down the EditorView, which reset the
      cursor to the document start and made typing appear right-to-left.
      
      * feat(web): add unsaved changes confirmation dialog in PlaygroundMessageEditor
      
      Implement a confirmation dialog to warn users about unsaved changes when attempting to leave the editor. This includes handling the beforeunload event to prevent accidental navigation away from the editor. Additionally, add tests to verify the dialog's behavior under various scenarios.
      
      * test(web): cover beforeunload guard and fade hydration suppression
      
      Address review feedback: add regression tests for the unsaved-changes
      beforeunload guard and the first-render fade suppression of hydrated
      content, and annotate getCachedMarkdown's return type.
      RedwindA committed
    • feat: channel test (#6917) · 4add708e
      * feat: channel test
      
      * fix: code smell
      Seefs committed