This is the multi-page printable view of this section. Click here to print.

Return to the regular view of this page.

Configuration Guide

Configuration guides for huatuo-bamai and huatuo-apiserver.

1 - huatuo-bamai Configuration

1. Overview

huatuo-bamai is the core collector of HUATUO (a BPF-based metrics and anomaly inspector). Its configuration file defines the data collection scope, probe enablement strategy, metric output format, anomaly detection rules, and logging behavior.

The configuration file uses TOML format and includes multiple sections such as global blacklist, logging, runtime resource limits, storage configuration, and AutoTracing. Each configuration item comes with detailed comments explaining its purpose, default value, and important notes. This document provides a clear and detailed English explanation for every configuration item to help users understand and safely customize the settings.

Note: Most parameters are provided as commented defaults (prefixed with #). Uncomment and adjust as needed. Changes take effect after restarting huatuo-bamai. In production, avoid enabling high-overhead features unnecessarily.

2. Global Blacklist

# Global tracing and metrics configuration.
#
# - BlackList
# Global blacklist for tracing and metrics.
#
BlackList = ["netdev_hw", "netdev_qdisc", "metax_gpu", "ascend_npu", "diskio", "tcp_retransmit", "mthreads_gpu"]
  • BlackList: Global blacklist for tracing and metrics.

    Modules or hardware to exclude from tracing and metric collection. The default is ["netdev_hw", "netdev_qdisc", "metax_gpu", "ascend_npu", "diskio", "tcp_retransmit", "mthreads_gpu"], which disables tracing and metrics for the network device hardware layer, qdisc statistics, Metax GPU, Ascend NPU, procfs-based disk I/O statistics, TCP retransmission tracing, and Moore Threads GPU. Remove diskio to enable disk I/O metrics, tcp_retransmit to enable TCP retransmission tracing, or mthreads_gpu to enable Moore Threads GPU metric collection on hosts with MT GPUs. Local correlation does not require the standalone dropwatch tracer. Supports arrays; extend as needed.

3. Logging

# Log Configuration
[Log]
    # - Level
    # The log level for huatuo-bamai: Debug, Info, Warn, Error, Panic.
    # Default: Info
    #
    # - File
    # Store logs to where the logging file is. If it is empty, don't write log
    # to any file.
    # Default: empty
    #
    # Level = "Info"
    # File = ""
  • Level: Log verbosity. Values: Debug, Info, Warn, Error, Panic. Default: Info. Use Info or Warn in production; Debug for troubleshooting.

  • File: Log file path.

    Specifies the path to the log file. If left empty, logs are not written to any file (output goes to stdout or system logs).

    Default: empty.

    Description: In containerized deployments, configure a specific path and integrate with a log collection system for persistence.

4. Runtime Resource Limits

Huatuo does not create its own cgroup by default. This section applies only when --enable-cgroup is passed; Kubernetes and systemd deployments should use their native resource controls.

# Runtime limits for the huatuo-bamai process.
[Runtime]
    # - StartupCPULimitCores
    # CPU limit during startup, in cores.
    # Default: 0.5
    #
    # - CPULimitCores
    # CPU limit after startup, in cores.
    # Default: 2.0
    #
    # - MemoryLimitMiB
    # Memory limit in MiB.
    # Default: 2048
    #
    # StartupCPULimitCores = 0.5
    # CPULimitCores = 2.0
    # MemoryLimitMiB = 2048
  • StartupCPULimitCores limits CPU usage during initialization. Default: 0.5 cores.
  • CPULimitCores limits CPU usage after startup. Default: 2.0 cores.
  • MemoryLimitMiB limits process memory. Default: 2048 MiB.

The configured values remain in their documented units. Memory is converted to bytes only when the cgroup limit is applied.

5. HTTP Server and On-demand Operations

# HTTP server configuration.
[HTTPServer]
    # - ListenAddress
    # Listen address in "host:port" form.
    # Default: ":19704"
    #
    # - MaxEventStreamClients
    # Maximum number of concurrent clients allowed to hold an open
    # /v1/events/watch SSE connection. Once the limit is reached, new requests
    # are rejected with HTTP 429 until an existing client disconnects.
    # Default: 100
    #
    # - EventStreamKeepAliveIntervalSeconds
    # Interval in seconds at which the server sends an SSE comment ping to each
    # connected client. The ping keeps the connection alive through load
    # balancers and proxies that would otherwise time out idle connections. If
    # writing the ping fails three consecutive times, the server closes the
    # connection.
    # Default: 30
    #
    # ListenAddress = ":19704"
    # MaxEventStreamClients = 100
    # EventStreamKeepAliveIntervalSeconds = 30

[HTTPServer.Auth]
    # Required service credential used by huatuo-apiserver.
    BearerToken = "REPLACE_WITH_RANDOM_HEX"

# Shared lifecycle policy for Profiling and Tracing operations.
[Operations]
    # MaxConcurrent = 10
    # LaunchTimeoutSeconds = 10
    # StopGracePeriodSeconds = 5
    # FinalizationTimeoutSeconds = 30
    # TerminalRetentionPeriodSeconds = 600

# Node-local profiler execution settings.
[Profiling]
    # AggregationIntervalSeconds = 10
    # MaxConcurrentProcesses = 10
    # CommandOutputLimitBytes = 65536
    # ToolDir = "/opt/huatuo/tools"
  • ListenAddress uses host:port form. An empty host listens on all interfaces.
  • HTTPServer.Auth.BearerToken is required and must match the independent Node credential configured for huatuo-apiserver. Replace the example value before deployment.
  • Operations.MaxConcurrent is one process-wide limit shared by Profiling and Tracing. New operations are rejected instead of queued when it is full.
  • The four operation time settings independently limit process launch, graceful stop, result finalization, and terminal-state retention.
  • Profiling.ToolDir is the shared external tool root, passed unchanged as profiler --tool-path. Java requires bin/asprof and lib/libasyncProfiler.so beneath this root; Python requires py-spy. Only the requested language’s tools are checked. Native profiling does not require this setting. Unsupported node environments reject the request without creating an operation.

The generated Node API exposes its contract at GET /openapi.json. Profiling and Tracing Start, Get, and Stop routes, POST /v1/events/watch, and PUT /v1/config require the service bearer token. /readyz, metrics, version, and the OpenAPI document remain public.

5.1 Update Configuration through the Node API

PUT /v1/config accepts one non-empty config object. Keys use the same dot-separated paths as the TOML structure, while values retain their JSON types:

curl -i -X PUT 'http://127.0.0.1:19704/v1/config' \
  -H 'Authorization: Bearer REPLACE_WITH_RANDOM_HEX' \
  -H 'Content-Type: application/json' \
  --data '{
    "config": {
      "BlackList": ["dropwatch", "netdev_hw"],
      "Runtime.CPULimitCores": 1.5,
      "Runtime.MemoryLimitMiB": 1024
    }
  }'

A successful update returns 204 No Content. The Node Agent validates the complete candidate configuration, atomically replaces the configuration file, and only then publishes the new in-memory snapshot. Validation or persistence failure leaves the current snapshot unchanged. Unknown keys, invalid value types, and empty update objects return 400 Bad Request.

Components that read configuration dynamically can observe the new snapshot without a restart. Settings consumed during startup, including the HTTP listener and authentication, storage initialization, and cgroup setup, are persisted but take effect only after restarting huatuo-bamai.

The event stream settings control POST /v1/events/watch. When MaxEventStreamClients is reached, new streams receive HTTP 429. EventStreamKeepAliveIntervalSeconds controls SSE heartbeat comments used to keep proxy and load-balancer connections alive. After three consecutive write failures, the server closes the stream. Set the interval below any upstream idle timeout; 15–60 seconds is typical.

6. Storage

6.1 Elasticsearch and OpenSearch Storage

# Storage configuration
[Storage]
    # Elasticsearch and OpenSearch Storage
    #
    # Disable ES/OS storage if one of Address, Username, Password is empty.
    # Store the tracing and events data of linux kernel to ES/OS.
    #
    # - Address
    # Port 9200 is commonly used for Elasticsearch/OpenSearch HTTP APIs.
    # e.g.
    # http://127.0.0.1:9200
    # https://127.0.0.1:9200
    #
    # - Index
    # Elasticsearch or OpenSearch index, a logical namespace that holds a collection of
    # documents for huatuo-bamai.
    # Default: huatuo_bamai
    #
    # - Username
    # - Password
    # Address, Username, and Password must be either all empty (disabled) or
    # all configured (enabled). Partial connection settings are invalid.
    #
    [Storage.Elasticsearch]
        # Address = "http://127.0.0.1:9200"
        # Index = "huatuo_bamai"
        # Username = "elastic"
        # Password = "REPLACE_WITH_PASSWORD"
  • Address: ElasticSearch/OpenSearch service address.

    No default value.

    Description: Used to store kernel tracing and event data. ES/OS storage is disabled when Address, Username, and Password are all empty. All three values are required when storage is enabled; a partial configuration prevents startup.

  • Index: Index name.

    Default: huatuo_bamai.

    Description: Logical namespace for organizing huatuo-bamai tracing and event documents.

  • Username: Authentication username.

    No default value.

    Description: Used for Basic Auth.

  • Password: Authentication password.

    No default value.

    Description: Used together with the username. In production, use a strong password and enable TLS encryption.

Overall: ES/OS storage persists kernel tracing and event data for later search and analysis.

6.2 Local File Storage

# LocalFile Storage
#
# Store data to local directory for troubleshooting on the host machine.
#
# - Path
# The directory for storing data. If the Path is empty, LocalFile will be disabled.
# Default: "huatuo-local"
#
# - RotationSizeMiB
# The maximum size in Megabytes of a record file before it gets rotated
# per kernel tracer.
# Default: 100MB
#
# - MaxRotatedFiles
# The maximum number of old log files to retain for per tracer.
# Default: 10
#
[Storage.LocalFile]
    # Path = "huatuo-local"
    # RotationSizeMiB = 100
    # MaxRotatedFiles = 10
  • Path: Local data storage directory.

    Default: huatuo-local. If empty, local file storage is disabled.

    Description: Stores data locally on the host for on-site troubleshooting. Use an absolute path.

  • RotationSizeMiB: Single file rotation size.

    Maximum size of a record file before rotation (per tracer).

    Default: 100 MB.

    Description: Prevents any single file from growing too large and consuming excessive disk space.

  • MaxRotatedFiles: Maximum number of rotated files to retain.

    Default: 10.

    Description: Oldest files are automatically deleted once the limit is reached, controlling disk usage.

7. Automatic Tracing

The automatic tracing module is one of HUATUO’s intelligent features. It triggers specific performance tracing based on thresholds, reducing manual intervention.

7.1 CPUIdle Automatic Tracing — Sudden High CPU Usage in Containers

# Autotracing configuration 
[AutoTracing]
    # cpuidle
    #
    # For sudden high CPU usage in containers.
    #
    # - UserThreshold
    # User CPU usage threshold, when cpu usage reaches this threshold, cpu
    # performance tracing will be triggered.
    # Default: 75%
    #
    # - SysThreshold
    # System CPU usage threshold, when reaching this threshold, cpu performance
    # tracing will be triggered.
    # Default: 45%
    #
    # - UsageThreshold
    # The total cpu usage (system + user cpu usage) threshold, when reaching
    # this threshold, cpu performance tracing will be triggered.
    # Default: 45%
    #
    # - DeltaUserThreshold
    # The range of this user cpu changes within a short period of time.
    # Default: 45%
    #
    # - DeltaSysThreshold
    # The range of this system cpu changes within a short period of time.
    # Default: 20%
    #
    # - DeltaUsageThreshold
    # The range of this cpu usage changes within a short period of time.
    # Default: 55%
    #
    # - Interval
    # The sample interval of the cpu usage for all containers.
    # Default: 10s
    #
    # - IntervalTracing
    # Time since last run. Avoid frequently executing this tracing to prevent
    # performance impact.
    # Default: 1800s
    #
    # - RunTracingToolTimeout
    # Execution timeout of this tracing tool (seconds).
    # Default: 10s
    # 
# NOTE:
# Profiling triggers when:
# 1. UserThreshold AND DeltaUserThreshold are exceeded, or
# 2. SysThreshold AND DeltaSysThreshold are exceeded, or
# 3. UsageThreshold AND DeltaUsageThreshold are exceeded
    #
    [AutoTracing.CPUIdle]
        # UserThreshold = 75
        # SysThreshold = 45
        # UsageThreshold = 90
        # DeltaUserThreshold = 45
        # DeltaSysThreshold = 20
        # DeltaUsageThreshold = 55
        # Interval = 10
        # IntervalTracing = 1800
        # RunTracingToolTimeout = 10
  • UserThreshold: User-mode CPU usage threshold (%).

    Default: 75%.

  • SysThreshold: System-mode CPU usage threshold (%).

    Default: 45%.

  • UsageThreshold: Total CPU usage threshold (%).

    Default: 90% (as shown in comments).

  • DeltaUserThreshold: Short-term user CPU change threshold (%).

    Default: 45%.

  • DeltaSysThreshold: Short-term system CPU change threshold (%).

    Default: 20%.

  • DeltaUsageThreshold: Short-term total CPU change threshold (%).

    Default: 55%.

  • Interval: CPU usage sampling interval (seconds).

    Default: 10s.

  • IntervalTracing: Minimum interval between runs (seconds).

    Default: 1800s (30 minutes).

  • RunTracingToolTimeout: Single tracing execution timeout (seconds).

    Default: 10s.

Trigger Logic: Tracing runs when any of the following is true:

  1. Both UserThreshold and DeltaUserThreshold are met, or
  2. Both SysThreshold and DeltaSysThreshold are met, or
  3. Both UsageThreshold and DeltaUsageThreshold are met.

Filter Container Filtering: Use Included/Excluded rule arrays to control monitoring scope.

    # Each rule contains Field (filter field) and Pattern (regex).
    # Field: container_host_namespace | container_hostname | container_qos
    #
    # [[AutoTracing.CPUIdle.Filter.Excluded]]
    #     Field = "container_qos"
    #     Pattern = "besteffort"
    # [[AutoTracing.CPUIdle.Filter.Included]]
    #     Field = "container_host_namespace"
    #     Pattern = "^application-"
  • Filter: Container filtering rules. Defined using [[double-bracket]] syntax with multiple rules, each containing Field (filter field) and Pattern (regex). Filtering logic:

    • No rules: monitor all containers
    • Excluded only: blacklist, skip matched containers
    • Included only: whitelist, only monitor matched containers
    • Both: must match Included AND not match Excluded

    Default: no rules, all containers monitored.

7.2 CPUSys Automatic Tracing — Sudden High System CPU on Host

# cpusys
#
# For sudden high system cpu usage on the host machine.
#
# - SysThreshold
# System CPU usage threshold, when reaching this threshold, cpu performance
# tracing will be triggered.
# Default: 45%
#
# - DeltaSysThreshold
# The range of system cpu changes within a short period of time.
# Default: 20%
#
# - Interval
# The sample interval of the cpu usage for host machine.
# Default: 10s
#
# - IntervalTracing
# Minimum time between profiling runs.
# Default: 1800s
#
# - RunTracingToolTimeout
# Execution timeout of this tracing tool (seconds).
# Default: 10s
#
# NOTE:
# Profiling triggers when:
# SysThreshold AND DeltaSysThreshold are exceeded.
#
[AutoTracing.CPUSys]
	# SysThreshold = 45
	# DeltaSysThreshold = 20
	# Interval = 10
	# IntervalTracing = 1800
	# RunTracingToolTimeout = 10
  • SysThreshold: System CPU usage threshold (%).

    Default: 45%.

  • DeltaSysThreshold: Short-term system CPU change threshold (%).

    Default: 20%.

  • Interval: Host CPU usage sampling interval (seconds).

    Default: 10s.

  • IntervalTracing: Minimum time between profiling runs. Default: 1800s.

  • RunTracingToolTimeout: Tracing execution timeout (seconds).

    Default: 10s.

Trigger Logic: Tracing is triggered when both SysThreshold and DeltaSysThreshold are satisfied.

7.3 Dload AutoTracing — D-State Task Profiling for Containers

# dload
#
# linux tasks D state profiling for containers.
#
# - ThresholdLoad
# Load average threshold. When exceeded, D-state profiling triggers.
# Default: 5
#
# - Interval
# The sample interval of the load for all containers.
# Default: 10s
#
# - IntervalTracing
# Time since last run. Avoid frequently executing this tracing to prevent
# performance impact.
# Default: 1800s
#
[AutoTracing.Dload]
	# ThresholdLoad = 5
	# Interval = 10
	# IntervalTracing = 1800
  • ThresholdLoad: System load average (loadavg) threshold for containers.

    Default: 5. Triggers D-state (uninterruptible sleep) task profiling when loadavg reaches this value.

  • Interval: Monitoring interval.

    Default: 10s.

  • IntervalTracing: Minimum time between consecutive tracings.

    Default: 1800s (30 minutes).

7.4 IOTracing AutoTracing — Container IO Performance Profiling

# iotracing
#
# io profiling for containers.
#
# - WbpsThreshold
# Max write bytes per second threshold. When exceeded, iotracing is triggered.
# For NVMe devices, UtilThreshold must also be met.
# Default: 1500 MB/s
#
# - RbpsThreshold
# Max read bytes per second threshold. When exceeded, iotracing is triggered.
# For NVMe devices, UtilThreshold must also be met.
# Default: 2000 MB/s
#
# - UtilThreshold
# Disk utilization (%). Consistently above 80-90% indicates a bottleneck.
# Default: 90%
#
# - AwaitThreshold
# Await (Average IO wait time in ms): High values indicate slow disk response times.
# Default: 100ms
#
# - RunTracingToolTimeout
# Execution timeout of this tracing tool (seconds).
# Default: 10s
#
# - MaxProcDump
# The number of processes displayed by iotracing tool.
# Default: 10
#
# - MaxFilesPerProcDump
# The number of files per process displayed by iotracing tool.
# Default: 5
#
[AutoTracing.IOTracing]
	# WbpsThreshold = 1500
	# RbpsThreshold = 2000
	# UtilThreshold = 90
	# AwaitThreshold = 100
	# RunTracingToolTimeout = 10
	# MaxProcDump = 10
	# MaxFilesPerProcDump = 5
  • WbpsThreshold: Max write bytes per second threshold (MB/s).

    Default: 1500. (For NVMe, must also meet UtilThreshold.)

  • RbpsThreshold: Max read bytes per second threshold (MB/s).

    Default: 2000.

  • UtilThreshold: Disk utilization threshold (%).

    Default: 90%.

  • AwaitThreshold: Average IO wait time threshold (ms).

    Default: 100ms.

  • RunIOTracingTimeout: IO tracing tool timeout (seconds).

    Default: 10s.

  • MaxProcDump: Maximum number of processes to display.

    Default: 10.

  • MaxFilesPerProcDump: Maximum files per process to display.

    Default: 5.

Description: Used for diagnosing IO hotspots in containers, especially under high disk load.

7.5 MemoryBurst AutoTracing

This module detects sudden memory usage spikes on the host and automatically captures kernel context to help diagnose memory pressure events.

# memory burst
#
# Capture kernel context on sudden host memory usage spikes.
#
# - Interval
# Memory usage sampling interval (seconds).
# Default: 10s
#
# - DeltaMemoryBurst
# Growth percentage threshold for memory usage. 100% means, e.g.,
# memory usage increased from 200MB to 400MB.
# Default: 100%
#
# - DeltaAnonThreshold
# Growth percentage threshold for anonymous memory. 100% means, e.g.,
# anon memory increased from 200MB to 400MB.
# Default: 70%
#
# - IntervalTracing
# Time since last run. Avoid frequently executing this tracing
# to prevent performance impact.
# Default: 1800s
#
# - DumpProcessMaxNum
# Number of processes to dump when triggered.
# Default: 10
#
[AutoTracing.MemoryBurst]
	# DeltaMemoryBurst = 100
	# DeltaAnonThreshold = 70
	# Interval = 10
	# IntervalTracing = 1800
	# SlidingWindowLength = 60
	# DumpProcessMaxNum = 10
  • DeltaMemoryBurst: Memory usage burst growth percentage threshold.

    Default: 100%.

  • DeltaAnonThreshold: Anonymous memory burst growth percentage threshold.

    Default: 70%.

  • Interval: Memory usage sampling interval (seconds).

    Default: 10s.

  • IntervalTracing: Minimum interval between runs (seconds).

    Default: 1800s.

  • SlidingWindowLength: Sliding window length (seconds).

    Default: 60s.

  • DumpProcessMaxNum: Maximum processes to dump on trigger.

    Default: 10.

7.6 Memory Threshold Runtime Snapshots

memory_threshold_snapshot is enabled by default. Add it to the global BlackList and restart huatuo-bamai to disable it. This feature attempts a Go, HotSpot, or CPython snapshot after a container memory-pressure notification; completion before OOM is not guaranteed. Candidates are ranked by an approximate kernel OOM score.

The persisted victim_pid, victim_process_name, and victim_oom_score_adj fields describe the process selected for snapshot capture.

[AutoTracing.MemoryThresholdSnapshot]
    # ThresholdPercent = 90
    # IntervalTracing = 300
    # RunTracingToolTimeout = 2
    # MaxMemoryObjectEntries = 10

Commented values are defaults.

Parameter Meaning
ThresholdPercent Required memory usage-to-limit percentage, from 1 to 100
IntervalTracing Node-wide minimum interval after a successful or failed capture attempt, in seconds; defaults to 300 and must be positive
RunTracingToolTimeout Cooperative capture timeout shared by Go, Java, and Python, in seconds; defaults to 2 and must be positive
MaxMemoryObjectEntries Maximum number of ranked memory object entries in a snapshot, from 1 to 100; defaults to 10; final JSON is trimmed to at most 512 KiB

Runtime detection and persistence are not included in the capture budget. Runtime detection has no separate timeout. Timeouts cannot interrupt synchronous reads already executing, so they do not bound the total operation time.

Trigger conditions:

  • Cgroup v1: Registers the memory threshold corresponding to ThresholdPercent through cgroup.event_control.
  • Cgroup v2: Watches increases in the high or max counters of memory.events.local (falling back to memory.events when absent), then checks memory.current / memory.max against the configured percentage. This feature does not set memory.high. When it is max, the hard-limit max counter can still trigger a check. This is a late notification and does not guarantee detection at the configured percentage or before OOM.

Both versions check usage once after registration and when the hard limit changes. One watcher manages all targets without periodic sampling. Repeated notifications for a target are coalesced; they do not count every crossing or report recovery below the threshold. Capture rechecks current usage and identity.

This feature requires the pod manager. It actively maintains the container view from shared CSS lifecycle hints and supplies an initial view followed by events with instance generations, init PIDs, and memory cgroup paths. The snapshot module no longer scans the cgroup tree. Subscription overflow recovers from a complete pod view; synchronization failure suspends capture and never implies deletion.

Running ordinary containers and restartable init sidecars are monitored; ephemeral debug containers and ordinary init containers are excluded. Ordinary registration errors and target capacity exhaustion are logged and the directory instance is skipped. Invalidated registrations are removed without automatic retries or backfilling. Repeated events, full views, and container generation changes do not restore a failed watch. A replacement directory, a new tracking lifetime after the previous container departs, or a huatuo-bamai restart allows a new attempt. Fatal errors such as host resource exhaustion still stop the current watcher. Each monitored container must have its own memory cgroup directory. A watch and its capture observations bind directly to one container instance. If registration returns a watch already owned by another container, the feature stops, cancels capture, and releases its watches. Correct the cgroup isolation and restart huatuo-bamai; the feature does not retry this conflict. Directory replacement invalidates the old registration through the watcher’s removal notification. Before capture and saving, the feature verifies the container generation and live binding, directory identity, and selected process identity and cgroup membership. Recorded memory usage and limits belong to that container’s cgroup.

See section 14 for deployment limitations and output lookup.

7.7 IRQTracing AutoTracing

This module detects abnormal irq+softirq utilization on one CPU and invokes irqtracing to collect softirq source and victim stacks.

[AutoTracing.IRQTracing]
    Interval = 2
    RunTracingToolTimeout = 3
    IntervalTracing = 300
    MaxEventsPerSecond = 1000
    MinCPUs = 3
    DeltaUsageThreshold = 20
    RelativeIncreaseThreshold = 30
    SustainedIntervals = 10
    UsageThreshold = 80
  • Interval: Sampling interval for per-CPU irq+softirq utilization from /proc/stat. Default: 2s.
  • RunTracingToolTimeout: Duration of one irqtracing collection. Default: 3s.
  • IntervalTracing: Minimum interval between triggers. Default: 300s.
  • MaxEventsPerSecond: Combined source and victim stack-sample limit per second on the traced CPU. Default: 1000. The daemon divides it as evenly as possible between softirq_raise and softirq_entry; the default is 500 events/s per stream. The value must be between 2 and 8589934590.
  • MinCPUs, DeltaUsageThreshold, and RelativeIncreaseThreshold: Configure the multi-CPU irq+softirq spike rule. The two thresholds are the increase in percentage points and the increase relative to the previous sample, respectively.
  • SustainedIntervals and UsageThreshold: Configure the consecutive sample count and utilization threshold for the single-CPU sustained rule.

7.8 Known Issue Filtering (IssuesList)

# Autotracing configuration.
#
# - IssuesList
# Known issue filters for autotracing.
#
[AutoTracing]
    IssuesList = []
  • IssuesList: Known issue filter. Format: [["name", "regex"], ...]. When a collected stack trace matches the regex, it is labeled with the issue name. Default [].

    Example: IssuesList = [["known_issue1", "softlockup"], ["known_issue2", "alloc_pages.*failed"]]

Note: Only supports dload tracing of known issues filtering, other events are not supported.

8. Event Tracing

This section captures key kernel events and latency, including scheduler tick intervals, memory reclaim, network receive latency, network device events, and packet drops. It is the core module for kernel-level anomaly context collection in HUATUO.

8.1 Scheduler Tick Interval Tracing

# linux kernel events capturing configuration
[EventTracing]
	# scheduler tick
	#
	# Trace long scheduler tick intervals.
	#
	# - IntervalThreshold
	# When the scheduler tick interval reaches the threshold, huatuo-bamai
	# will collect kernel context.
	# Default: 10000000 in nanoseconds, 10ms
	#
	[EventTracing.SchedTick]
		# IntervalThreshold = 10000000
  • IntervalThreshold: Scheduler tick interval threshold, in nanoseconds.

    Default: 10,000,000 ns (10ms).

    Description: The event infers CPU stalls from long scheduler tick intervals. It does not by itself prove that softirqs were disabled.

8.2 Memory Reclaim Blocking Tracing

# memreclaim
#
# The memory reclaim may block the process, if one process is blocked
# for a long time, reporting the events to userspace.
#
# - BlockedThreshold
# The blocked time when memory reclaiming.
# Default: 900000000ns, 900ms
#
[EventTracing.MemoryReclaim]
	# BlockedThreshold = 900000000
  • BlockedThreshold: Memory reclaim blocking time threshold (nanoseconds).

    Default: 900,000,000 ns (900ms). When a process is blocked by memory reclaim for longer than this time, an event is reported to userspace with context.

    Description: Memory reclaim blocking is a common cause of process stalls, especially in memory-constrained cloud-native environments.

8.3 Network Receive Latency Tracing

# networking rx latency
#
# linux net stack rx latency for every tcp skbs.
#
# - Driver2NetRx
# The latency from driver to net rx, e.g., netif_receive_skb.
# Default: 5ms
#
# - Driver2TCP
# The latency from driver to tcp rx, e.g., tcp_v4_rcv.
# Default: 10ms
#
# - Driver2Userspace
# The latency from driver to userspace copy data, e.g., skb_copy_datagram_iovec.
# Default: 115ms
#
# - ExcludedContainerQos
# Blacklist: skip containers whose qos level matches.
# Values: "guaranteed", "burstable", "besteffort" (case-insensitive).
# Default: [].
#
# - ExcludedHostNetnamespace
# Exclude packets in the host network namespace.
# Default: true
#
[EventTracing.NetRxLatency]
	# Driver2NetRx = 5
	# Driver2TCP = 10
	# Driver2Userspace = 115
	# ExcludedContainerQos = []
	ExcludedContainerQos = ["besteffort"]
	# ExcludedHostNetnamespace = true
  • Driver2NetRx: Latency threshold from driver to network receive layer (e.g., netif_receive_skb).

    Default: 5ms.

  • Driver2TCP: Latency threshold from driver to TCP receive (e.g., tcp_v4_rcv).

    Default: 10ms.

  • Driver2Userspace: Latency threshold from driver to userspace data copy (e.g., skb_copy_datagram_iovec).

    Default: 115ms.

  • ExcludedContainerQos: Container QoS levels to exclude (blacklist).

    Default: []. Corresponds to Kubernetes Pod QoS levels (Guaranteed, Burstable, BestEffort).

  • ExcludedHostNetnamespace: Whether to exclude packets in the host network namespace.

    Default: true.

8.4 Network Device Event Monitoring

# netdev events
#
# Monitor network device events.
#
# - DeviceList
# The net devices we monitor.
# Default: [] (empty, meaning no devices).
#
[EventTracing.Netdev]
	DeviceList = ["eth0", "eth1", "bond4", "lo"]
  • DeviceList: List of network device full-match regex patterns to monitor. Literal names such as "eth0" keep exact-match behavior; patterns such as "bond[0-9]+" can select multiple devices.

    Default example includes “eth0”, “eth1”, “bond4”, “lo”. An empty list means no devices are monitored.

    Description: Monitors physical link status events for specified network interfaces.

8.5 Packet Drop Monitoring

[EventTracing.Dropwatch]
    # Filter for standalone dropwatch.
    # Default: "tcp"
    Filter = "tcp"

    # Forwarded to dropwatch --max-events-per-second.
    # Default: 100; 0 disables rate limiting.
    MaxEventsPerSecond = 100

    # Reserved configuration field. It is not currently consumed by the
    # dropwatch event path and therefore has no filtering effect.
    # Default: []
    ExcludeContainers = []
  • Filter: tcpdump-style packet filter passed only to standalone dropwatch. TCP retransmission correlation uses TCPRetransmit.Filter for both of its inputs.

    Default: "tcp".

  • MaxEventsPerSecond: Maximum number of dropwatch events emitted by BPF per second.

    Default: 100. Set to 0 to disable rate limiting.

  • ExcludeContainers: Reserved container-exclusion list.

    Default: []. The field exists in the configuration schema, but the current dropwatch event path does not read or forward it, so configuring it has no effect. Use EventTracing.IssuesList for operator-defined dropwatch call-stack suppression.

8.6 TCP Retransmission Tracing ([EventTracing.TCPRetransmit])

[EventTracing.TCPRetransmit]
    # Retransmission filter. Local correlation applies it to both inputs.
    # Default: empty (no flag when disabled; "tcp" when enabled).
    Filter = ""

    # Forwarded as tcpshark --enable-tlp. Default: false.
    EnableTLP = false

    # Run tcpshark with an embedded dropwatch source. Default: false.
    EnableDropwatch = false

    # Forwarded as tcpshark --max-events-per-second.
    # Default: 100; 0 disables rate limiting.
    MaxEventsPerSecond = 100
  • EnableTLP: Whether to collect tcp_send_loss_probe events.

    Default: false.

  • Filter: Tcpdump-style retransmission filter used in both modes. Local correlation applies the normalized expression to both tcpshark inputs and defaults an empty value to tcp. When correlation is disabled, an empty value passes no --filter flag. Dropwatch.Filter independently controls standalone dropwatch.

  • EnableDropwatch: Whether tcpshark should load a private dropwatch source and finalize retransmissions locally. The default is false. tcp_retransmit must be removed from BlackList; standalone dropwatch may remain blacklisted. Retransmissions wait up to 100 ms for delayed delivery, and candidate drops must precede them by no more than one second in kernel monotonic time. The embedded source automatically detects and enables devlink DROP traps. A strict same-namespace match reports software or hardware according to source, with the same drop_source, drop_reason, and hardware drop_reason_group semantics as dropwatch; every finalized event has one correlation_reason (matched, unsupported, warmup, wait_timeout, queue_full, or interrupted). warmup applies only to expired waits whose retransmit timestamp predates source readiness; retransmissions at or after readiness receive wait_timeout on expiry, and other termination paths retain their own reasons. Unmatched events report drop_location=unknown with the namespace diagnostic and dropwatch counters, including map-counter availability.

  • MaxEventsPerSecond: Maximum TCP retransmission events emitted by BPF per second. Correlation mode gives embedded dropwatch an independent limiter with the same value, so 100 permits up to 100 events/s on each input.

    Default: 100. Set to 0 for unlimited output. When the limit is exceeded, tcpshark logs rate limit hit.

8.7 Hardware Error Event Tracing (EventTracing.Ras)

# ras
#
# Hardware error event tracing (RAS: Reliability, Availability, Serviceability).
# Captures MCE, EDAC, ACPI/GHES, PCIe AER, and MCE threshold (THR) events via eBPF.
#
# - MceThrBackoff
# Minimum interval in seconds between consecutive MCE threshold (THR) event saves.
# THR events are fired by the local-APIC threshold interrupt and can storm at high
# frequency; this cooldown prevents flooding storage with redundant records.
# Default: 1800s (30 minutes)
#
[EventTracing.Ras]
    # MceThrBackoff = 1800
  • MceThrBackoff: Minimum cooldown in seconds between MCE threshold (THR) event saves.

    Default: 1800s (30 minutes).

    Description: THR events are generated by the CPU’s local-APIC threshold interrupt when correctable hardware errors accumulate. These can fire at very high frequency during hardware degradation. The backoff suppresses redundant saves while ensuring at least one record is captured per interval. Lower values provide more granular event records at the cost of higher storage throughput; in environments with frequent correctable errors, consider raising this value to reduce noise.

8.8 Moore Threads GPU Event Tracing (EventTracing.MthreadsGPU)

# mthreads_gpu
#
# Moore Threads GPU XID error event tracing.
[EventTracing.MthreadsGPU]
    # MthreadsXidLevel = ""
  • MthreadsXidLevel: Minimum severity level for XID error reporting.

    Valid values: "" (disabled), "notify", "warning", "fatal".

    Default: "" (disabled).

    Description: Controls which XID error events are reported. Set to enable XID error tracking for Moore Threads GPUs. XID errors below the specified severity level are filtered out. To enable this feature, ensure that mthreads_xid is removed from the global BlackList (if present) and set this field to one of the valid severity levels ("notify", "warning", "fatal"). Additionally, the MUSA driver must be installed on the host. The feature polls /proc/driver/musa/gpu*/event_report files every second to capture XID error events. Each XID event contains detailed information including UUID, XID ID, severity, scope, PCI BDF, process ID, and additional context.

8.9 Known Issue Filtering (IssuesList)

# Linux kernel event tracing configuration.
#
# - IssuesList
# Known issue filters for event tracing.
#
[EventTracing]
    IssuesList = []
  • IssuesList: Known-issue suppression rules in the form [["name", "regex"], ...]. Default [].

    For net_rx_latency, each regex is matched against the generated event title. For dropwatch, it is matched against the newline-joined kernel call stack. A match causes the event to be discarded; the configured name identifies the rule but is not added to the saved event.

    Example: IssuesList = [["ignored_process", "comm=ignored_process"], ["neighbor_cleanup", "neigh_invalidate/"]]

9. Metric Collector

This section defines collection rules for various system and network metrics. All Included/Excluded fields share the same filter logic (regex):

  • No rules: all items are collected
  • Excluded only: blacklist, matched items are skipped
  • Included only: whitelist, only matched items are collected
  • Both: must match Included AND not match Excluded

9.1 Netdev Statistics

# Metric Collector
[MetricCollector]
	# Netdev statistic
	#
	# - EnableNetlink
	# Use netlink instead of procfs net/dev to get netdev statistic.
	# Only support the host environment to use `netlink` now.
	# Default is "false".
	#
	# - DeviceIncluded
	# Accept special devices in netdev statistic.
	# Default: "" (empty), meaning include all.
	#
	# - DeviceExcluded
	# Exclude special devices in netdev statistic.
	# Default: "" (empty), meaning exclude nothing.
	#
	# Filter logic see MetricCollector section header.
	#
	[MetricCollector.NetdevStats]
		# EnableNetlink = false
		# DeviceIncluded = ""
		DeviceExcluded = "^(lo)|(docker\\w*)|(veth\\w*)$"
  • EnableNetlink: Use netlink instead of procfs to collect netdev statistics.

    Default: false. Currently only supported on the host.

  • DeviceIncluded: Regex to include specific devices. Default: include all.

  • DeviceExcluded: Regex to exclude devices. Example: “^(lo)|(docker\w*)|(veth\w*)$”, meaning exclude loopback, docker, and veth interfaces.

9.2 Netdev DCB Collection

# netdev dcb, DCB (Data Center Bridging)
#
# Collecting the DCB PFC (Priority-based Flow Control).
#
# - DeviceList
# The net devices we monitor.
# Default: [] (empty, meaning no devices).
#
[MetricCollector.NetdevDCB]
	DeviceList = ["eth0", "eth1"]
  • DeviceList: List of network device full-match regex patterns for which DCB (Data Center Bridging) PFC information is collected.

    Default: empty.

9.3 Netdev Hardware Statistics

# netdev hardware statistic
#
# Collecting the hardware statistic of net devices, e.g, rx_dropped.
#
# - DeviceList
# The net devices we monitor.
# Default: [] (empty, meaning no devices).
#
[MetricCollector.NetdevHW]
	DeviceList = ["eth0", "eth1"]
  • DeviceList: List of network device full-match regex patterns for hardware-level statistics (e.g., rx_dropped).

    Default: empty.

9.4 Qdisc Collection

# Qdisc
#
# - DeviceIncluded / DeviceExcluded
# Same as above.
#
[MetricCollector.Qdisc]
	# DeviceIncluded = ""
	DeviceExcluded = "^(lo)|(docker\\w*)|(veth\\w*)$"
  • DeviceIncluded / DeviceExcluded: Same as above.

9.5 vmstat Metric Collection

# vmstat
#
# This metric supports host vmstat and cgroup vmstat.
# - IncludedOnHost / ExcludedOnHost: same as above, for host /proc/vmstat.
# - IncludedOnContainer / ExcludedOnContainer: same, for cgroup containers memory.stat.
#
[MetricCollector.Vmstat]
	IncludedOnHost = "allocstall|nr_active_anon|nr_active_file|nr_boost_pages|nr_dirty|nr_free_pages|nr_inactive_anon|nr_inactive_file|nr_kswapd_boost|nr_mlock|nr_shmem|nr_slab_reclaimable|nr_slab_unreclaimable|nr_unevictable|nr_writeback|numa_pages_migrated|pgdeactivate|pgrefill|pgscan_direct|pgscan_kswapd|pgsteal_direct|pgsteal_kswapd"
	ExcludedOnHost = "total"
	IncludedOnContainer = "active_anon|active_file|dirty|inactive_anon|inactive_file|pgdeactivate|pgrefill|pgscan_direct|pgscan_kswapd|pgsteal_direct|pgsteal_kswapd|shmem|unevictable|writeback|pgscan_globaldirect|pgscan_globalkswapd|pgscan_cswapd|pgsteal_cswapd|pgsteal_globaldirect|pgsteal_globalkswapd"
	ExcludedOnContainer = "total"
  • IncludedOnHost / ExcludedOnHost: Filter fields for host /proc/vmstat.

  • IncludedOnContainer / ExcludedOnContainer: Filter fields for container cgroup memory.stat.

9.6 Other Metric Collections

# MemoryEvents/Netstat/MountPointStat
#
# - Included / Excluded: same as above.
# - MountPointsIncluded: whitelist only (no Excluded), same logic.
#
[MetricCollector.MemoryEvents]
	Included = "watermark_inc|watermark_dec"
	# Excluded = ""
[MetricCollector.Netstat]
	# Excluded = ""
	# Included = ""

# MountPointStat
[MetricCollector.MountPointStat]
	MountPointsIncluded = "(^/home$)|(^/$)|(^/boot$)"
  • Included / Excluded: Same as above.

  • MountPointsIncluded: Regex for mount points to collect. Default includes /, /home, /boot.

9.7 Moore Threads GPU Metrics

# MetricCollector.Mthreads
#
# Moore Threads GPU metric collection via the MTML (Moore Threads Management
# Library) shared library. The library is discovered automatically at startup
# using SONAME search (libmtml.so.2, then libmtml.so) through the system
# dynamic linker; no hardcoded path is required.
#
# Remove "mthreads_gpu" from BlackList to enable this collector.
#
# - EnableHealth
# Enable health metrics: temperature, power, utilization, clocks, fans, pstate, VPU.
# Default: true
#
# - EnablePCIe
# Enable PCIe link metrics: current speed/width and replay counter.
# Default: false
#
# - EnableMTLink
# Enable MtLink interconnect metrics: per-link state and bandwidth.
# Default: false
#
[MetricCollector.Mthreads]
    # EnableHealth = true
    # EnablePCIe = false
    # EnableMTLink = false
  • EnableHealth: Controls collection of health-related metrics.

    Default: true. When enabled, collects: GPU/memory temperature (gpu_temperature_celsius, memory_temperature_celsius), power usage and limits (device_power_watts, gpu_power_limit_watts, gpu_power_default_limit_watts), GPU/memory utilization (gpu_utilization_percent, memory_utilization_percent), clock frequencies (gpu_clock_mhz, gpu_max_clock_mhz, memory_clock_mhz, memory_max_clock_mhz), voltage (gpu_voltage_volts), memory capacity (memory_total_bytes, memory_used_bytes), fan speed (fan_rpm, fan_speed_percent), performance state (gpu_pstate), and VPU metrics (vpu_utilization_percent, vpu_encoder_utilization_percent, vpu_decoder_utilization_percent, vpu_clock_mhz).

  • EnablePCIe: Controls collection of PCIe link metrics.

    Default: false. When enabled, collects: current PCIe link speed and width (pcie_link_speed_gt_per_sec, pcie_link_width_lanes), max capability (pcie_link_max_speed_gt_per_sec, pcie_link_max_width_lanes), and replay counter (pcie_replay_total).

  • EnableMTLink: Controls collection of MtLink interconnect metrics.

    Default: false. When enabled, collects: device-level static specs (per-link bandwidth mtlink_link_bandwidth_gb_s and link count mtlink_link_count) and per-link state (mtlink_state).

Library discovery: At startup, the collector searches for libmtml.so.2 then libmtml.so via the system dynamic linker (respecting LD_LIBRARY_PATH and /etc/ld.so.cache). If no library is found, the collector logs a warning and is disabled for the lifetime of the process. Library discovery is performed only at startup: changing LD_LIBRARY_PATH or installing a new MTML version also requires a restart.

Hot-reload semantics: EnableHealth, EnablePCIe, and EnableMTLink are read from the latest config snapshot on every scrape, so toggling them takes effect on the next Prometheus scrape without restarting huatuo-bamai. A false → true → false transition emits and suppresses the corresponding metric groups on the next scrape after each change.

Note: enabling the collector itself (i.e. removing mthreads_gpu from BlackList after the process has already started with the collector disabled because libmtml.so was missing at startup) requires a restart. The collector factory runs only during initialization, so a successful late library load will not register a new collector.

10. Pod

This section configures how to fetch Pod information from kubelet to enable container/Pod-level labeling and metric isolation.

# Pod Configuration
#
# Configure these parameters for fetching pods from kubelet.
#
# - KubeletReadOnlyPort
# The KubeletReadOnlyPort is kubelet read-only port for the Kubelet to serve on with
# no authentication/authorization. The port number must be between 1 and 65535, inclusive.
# Setting this field to 0 disables fetching pods from kubelet read-only service.
# Default: 10255
#
# - KubeletAuthorizedPort
# The port is the HTTPs port of the kubelet. The port number must be between 1 and 65535,
# inclusive. Setting this field to 0 disables fetching pods from kubelet HTTPS port.
# Default: 10250
#
# - KubeletClientCertPath
# https://kubernetes.io/docs/setup/best-practices/certificates/
#
# Client certificate and private key file name. One file or two files:
# "/path/to/xxx-kubelet-client.crt,/path/to/xxx-kubelet-client.key",
# "/path/to/kubelet-client-current.pem"
#
# You can disable this kubelet fetching pods, for bare metal service, by
# KubeletReadOnlyPort = 0, and KubeletAuthorizedPort = 0.
#
[Pod]
	KubeletClientCertPath = "/etc/kubernetes/pki/apiserver-kubelet-client.crt,/etc/kubernetes/pki/apiserver-kubelet-client.key"
  • KubeletReadOnlyPort: Kubelet read-only port.

    Default: 10255. Set to 0 to disable this method.

  • KubeletAuthorizedPort: Kubelet HTTPS authorized port.

    Default: 10250. Set to 0 to disable.

  • KubeletClientCertPath: Path to kubelet client certificate and private key. Supports comma-separated files or single PEM file.

    Description: Used for mTLS authentication on the HTTPS port. In non-Kubernetes (bare-metal) environments, set both ports to 0 to disable Pod fetching.

11. CLI Flags

huatuo-bamai supports the following command-line flags:

huatuo-bamai --region <region> [options]
Flag Description Default
--config Configuration file name huatuo-bamai.conf
--config-dir Configuration file directory conf
--bpf-dir BPF object file directory bpf
--tools-bin-dir Tracing tool binary directory bin
--region Deployment region (required) -
--disable-kubelet Disable kubelet Pod fetching false
--disable-storage Disable storage backends false
--enable-cgroup Enable self cgroup resource limits (disabled by default) false
--disable-tracing Disable specified tracing modules (may be repeated) -
--log-debug Force log level to Debug false
--dry-run Load-only test; exit gracefully after startup false
--procfs-prefix procfs mount point prefix -

12. Configuration Override Precedence

When the same configuration item is set in both command-line flags and the configuration file, the following precedence applies:

CLI flag > Configuration file > Built-in default

Specific rules:

  1. Log level: --log-debug > config file [Log] Level > built-in default Info

    • --log-debug has the highest priority and forces the log level to Debug regardless of the Level value in the configuration file.
    • An explicit Level in the configuration file overrides the built-in default.
    • If neither is set, the default Info is used.
  2. Tracing blacklist: --disable-tracing is merged with the configuration file BlackList (they complement each other rather than override).

  3. Other boolean switches (--disable-kubelet, --disable-storage): When explicitly set on the command line, they override the configuration file.

13. Best Practices and Important Notes

  • Resource Control: Kubernetes uses Pod resources and systemd uses service limits. Use --enable-cgroup and [Runtime] only for direct execution without an external manager.
  • Storage Choice: For small-scale deployments, prefer [Storage.LocalFile] for local troubleshooting. For large clusters, configure Elasticsearch for centralized storage and querying.
  • AutoTracing Tuning: Adjust thresholds based on workload characteristics. Thresholds that are too low cause frequent triggering; thresholds that are too high may miss issues. Validate gradually in a test environment.
  • Security: Use strong passwords for ES configuration and consider enabling HTTPS. Avoid hard-coding sensitive information in the configuration file.
  • Compatibility: Configuration parameters may be affected by kernel version and hardware environment. Always verify with the official HUATUO documentation for your specific setup.

By properly configuring huatuo-bamai.conf, you can fully leverage HUATUO’s capabilities in kernel-level anomaly detection and intelligent tracing, significantly improving observability and troubleshooting efficiency in cloud-native systems.

If you need deeper customization for a specific scenario, feel free to provide more details about your environment.

14. Memory Threshold Snapshot Deployment and Troubleshooting

14.1 Requirements and Limitations

  • Requires Linux, memory cgroup v1/v2, host PID/procfs/cgroup views, kubelet metadata, kernel BTF, and BPF load/attach permissions.
  • Requires target-memory read permission (usually CAP_SYS_PTRACE), procfs/cgroup access, and v1 cgroup.event_control write access. Yama, SELinux, or AppArmor may block access.
  • Selects only direct cgroup members, excluding oom_score_adj = -1000. Selection is skipped above 4096 PIDs, 64 KiB of PID data, or a one-second budget.
  • At most 4096 containers are watched. Targets come from pod events without scanning the cgroup tree. Ordinary registration errors and capacity exhaustion are logged and skipped; invalidated registrations are removed without retries or backfilling. Resource recovery or a full view does not re-register the same instance, so pressure monitoring may remain unavailable for that instance.
  • Failed identity checks or missing container metadata prevent capture or saving.
  • Process selection and capture both use /proc in Huatuo’s PID namespace. Changing --procfs-prefix does not redirect memory snapshot reads.

This table describes experimental implementation coverage, not validation of every listed version:

Runtime Experimental coverage Main limitations
Go Go 1.18–1.26, 64-bit ELF Instruction recovery for stripped executables is x86-64 only
Java Java 8+, little-endian 64-bit ELF HotSpot, G1 GC Requires recognizable VMStruct/VMType metadata
Python CPython 3.8–3.14, little-endian 64-bit ELF Requires discoverable _PyRuntime and recognizable version/layout

Recorded manual validation: x86-64 Linux, cgroup v1 (legacy/hybrid), Go 1.24.0. Verify pressure-triggered, non-empty snapshots in your environment; skipped tests and unavailable results do not prove compatibility.

14.2 Output and Troubleshooting

Info logs record watcher state and capture attempts. Process selection and persistence details use Debug logs. Locate attempts by container/cgroup and use process-selection logs to identify the PID. Inspect stored snapshot status and status_reason fields for runtime diagnostics. If a capture attempt starts but does not finish, inspect bamai’s /debug/pprof/goroutine?debug=2 with appropriate authorization for the blocked stack. No logs alone do not prove that monitoring has stopped.

tracer_data.process_memory reads /proc/<pid>/status once after runtime detection returns without error, including for unrecognized runtimes such as C/C++. Provider failures produce a failed snapshot and retain this summary. Detection or output-processing errors prevent persistence; inspect the capture attempt logs for the error. Target identity changes or cancellation discard the result. It provides no PSS, mapping rankings, or allocation stacks.

Field (bytes) Source / meaning
virtual_bytes VmSize, virtual address space, not physical memory usage
rss_bytes VmRSS, resident memory
rss_anon_bytes RssAnon, anonymous resident memory
rss_file_bytes RssFile, file-backed resident memory
rss_shmem_bytes RssShmem, shared resident memory
swap_bytes VmSwap, swapped private anonymous memory, excluding shmem swap
page_table_bytes VmPTE, page-table memory

Missing/invalid fields are omitted, not zero-filled: status is partial, or unavailable with status_reason if nothing can be read. Values are approximate, not an OOM-time snapshot or proof of a leak. The selected process may not be the eventual OOM victim.

Results use the existing [Storage] configuration (section 6); no separate storage setup is needed. The LocalFile filename is memory_threshold_snapshot. Query tracing_documents with tracer_name = memory_threshold_snapshot and tracer_type = autotracing.

started_timestamp records when the capture attempt starts, before process selection. observed_timestamp records when snapshot collection starts, after process selection and before runtime detection.

The kind and measurement semantics of tracer_data.snapshot.entries are:

Runtime kind objects bytes
Go inuse_space_objects Estimated unfreed object count from published statistics, corrected for sampling and grouped by complete allocation stack Estimated unfreed bytes for those objects
Java object_class Estimated instance count grouped by class Estimated memory occupied by the objects themselves (shallow heap)
Python gc_tracked_object_type GC-tracked object count grouped by type Estimated shallow size of those objects

Go’s bytes and objects correspond to pprof’s inuse_space and inuse_objects metrics. Java does not compute retained heap; Python covers only GC-tracked objects, not the entire Python heap. Older records may use allocation_site for Go and object_type for Java; consumers reading historical records should accept those values.

For Go entries, name remains the plain function name of the first non-runtime frame, or the first frame when all frames belong to the runtime. stack remains a string array ordered from the allocation site toward its callers. When both file and line are available, each frame uses function, source_file:line, with one space after the comma. For example:

{
  "kind": "inuse_space_objects",
  "name": "example/cache.allocate",
  "bytes": 131072,
  "objects": 32,
  "average_bytes": 4096,
  "stack": [
    "example/cache.allocate, example/cache/cache.go:123",
    "example/service.load, example/service/load.go:58",
    "main.main, example/cmd/server/main.go:42"
  ]
}

Locations come from the binary’s Go line table; source files need not exist on the target machine. Paths retain their compiled values, which may reflect -trimpath and may not identify files on the target machine. A missing file or an invalid line leaves only the function name; an unresolved function leaves the hexadecimal PC. Consumers should also accept historical frames containing only names or addresses. Function names and paths may themselves contain commas, so splitting indiscriminately on commas is not reliable. Source locations identify allocation sites, not object owners, and do not guarantee a fully expanded inline call chain.

Go subtracts frees from allocations using only the runtime’s published active counters, then corrects for sampling. Unpublished future[0..2] counters are excluded. Publication is delayed until the corresponding GC sweep frees are accounted for, so recent allocation spikes may be absent. Results are external samples of the published heap profile, not heap usage at the trigger instant. Collection does not acquire runtime locks, flush counters, or trigger GC in the target. Fields or buckets may still reflect different publication stages, so the result does not provide pprof’s atomic consistency.

If a complete scan observes no nonzero active counters, it returns unavailable with status_reason set to Go heap profile has no published statistics, without falling back to future. This can occur before the first GC and does not mean that the process has no heap objects. Published counters whose sampled objects have all been freed can still produce complete with empty entries. An interrupted scan preserves partial and its reason rather than treating unobserved data as unpublished.

Go aggregates complete stack keys up to 32 frames for Go 1.18–1.22 and 1024 frames for Go 1.23–1.26; the shared output limit may shorten displayed stacks to 64 frames and sets output_truncated. An invalid bucket type, an excessive stack depth, or a cyclic bucket chain stops the scan with partial; repeated buckets are never counted twice.

Once a scan becomes partial, it stops traversing further buckets and selects at most MaxMemoryObjectEntries entries by descending byte count from the valid samples retained from earlier batches and the current batch, within the aggregation budget. The status_reason records only the first cause; finishing the current batch does not append further causes.

Any bucket header, record, or stack read failure, including a short read, fails the entire Go collection attempt and discards all runtime entries, including those from earlier batches. The collector reports failed with the read error; individual ranges are not retried.

Samples with a stack depth of zero produce no allocation-site entry; an empty stack alone does not make the scan partial.

Go snapshots require a known, enabled sampling rate and a nonempty bucket list. An unknown or disabled rate, or an empty bucket list, yields unavailable with a status_reason and no entries.

Go collection uses a single request timeout across runtime reads, scanning, ranking, and entry construction. When it expires, runtime entries are discarded and the collector emits failed with a timeout reason; it still attempts to read the process memory summary. Cancellation is cooperative, so an in-flight system call or non-cancelable parsing step can finish after the deadline.

Inspect tracer_data.snapshot.status (complete, partial, unavailable, or failed) together with status_reason, runtime_version, duration_ms, and output_truncated.

status_reason explains why collection is partial, unavailable, or failed; it is omitted when empty. Both snapshot and process_memory use this field. Older records may use reason; readers of historical and new records should accept both names.

duration_ms measures the provider stage, rounded up to milliseconds, for both successful and failed snapshots, including timeouts. It excludes runtime detection, process memory summary reads, output processing, and persistence.

Problem Checks
No output Enable and restart; check BlackList; see section 7.6 for v2 trigger conditions
Event without a candidate Check direct cgroup membership, OOM-kill eligibility, and enumeration limits
unavailable / failed Check runtime/layout restrictions, access permissions, container metadata, and target exit; inspect status_reason
Event stops after resource exhaustion Check RLIMIT_NOFILE, fs.inotify.max_user_watches, and fs.inotify.max_user_instances; adjust and restart; this stop does not stop other events

2 - huatuo-apiserver Configuration

1. Overview

huatuo-apiserver uses a strictly decoded TOML configuration file. Unknown or obsolete options prevent startup. Commented options use the built-in defaults shown below.

2. Logging and Runtime Limits

huatuo-apiserver does not create its own cgroup by default. The [Runtime] section applies only when --enable-cgroup is explicitly passed; Kubernetes and systemd deployments should use their native resource controls.

# Log Configuration
[Log]
    # - Level
    # The log level for huatuo-apiserver: Debug, Info, Warn, Error, Panic.
    # Default: Info
    #
    # Level = "Info"

# Runtime limits for the huatuo-apiserver process.
[Runtime]
    # - CPULimitCores
    # CPU limit in cores.
    # Default: 20
    #
    # - MemoryLimitMiB
    # Memory limit in MiB.
    # Default: 4096
    #
    # CPULimitCores = 20
    # MemoryLimitMiB = 4096
  • Log.Level supports Debug, Info, Warn, Error, and Panic.
  • CPULimitCores limits the API server process in CPU cores.
  • MemoryLimitMiB limits the API server process in MiB.

All resource limits must be greater than zero. --log-debug overrides Log.Level.

3. HTTP Server

# HTTP server configuration.
[APIServer]
    # - ListenAddress
    # Listen address in "host:port" form.
    # Default: ":12740"
    #
    # ListenAddress = ":12740"

    # Request rate limiting.
    [APIServer.RateLimit]
        # - RequestsPerSecond
        # Maximum process-wide request rate per second.
        # Default: 200
        #
        # - Burst
        # Maximum process-wide request burst.
        # Default: 200
        #
        # RequestsPerSecond = 200
        # Burst = 200

ListenAddress uses host:port form. An empty host listens on all interfaces. RateLimit is a process-wide token bucket; both values must be positive. HTTP timeouts and request-size limits are fixed service safeguards and are not user configurable.

4. Jobs and Agent Communication

# Job persistence.
[Jobs]
    # - StoreDSN
    # SQLite DSN. Relative paths are resolved from this file's directory.
    # Default: "jobs.db"
    #
    # StoreDSN = "jobs.db"

    # Job records use an internal revision for concurrent updates. Only records
    # written by this server version are supported; stop the old service and
    # remove or replace an incompatible database before starting this version.

    # Profiling and tracing retain independent quotas because their resource
    # costs and expected concurrency differ.
    [Jobs.Profiling]
        # - MaxConcurrentPerHost
        # Maximum concurrent profiling jobs on one host.
        # Default: 3
        #
        # - MaxConcurrent
        # Maximum concurrent profiling jobs across all hosts.
        # Default: 500
        #
        # MaxConcurrentPerHost = 3
        # MaxConcurrent = 500

    [Jobs.Tracing]
        # - MaxConcurrentPerHost
        # Maximum concurrent tracing jobs on one host.
        # Default: 5
        #
        # - MaxConcurrent
        # Maximum concurrent tracing jobs across all hosts.
        # Default: 1000
        #
        # MaxConcurrentPerHost = 5
        # MaxConcurrent = 1000

    # Apiserver-owned Job lifecycle policy.
    [Jobs.Controller]
        # StatusPollIntervalSeconds = 5
        # PendingTimeoutSeconds = 30
        # CompletionGracePeriodSeconds = 60
        # NodeUnavailableGracePeriodSeconds = 30
        # JobRetentionPeriodHours = 720

# huatuo-bamai Agent HTTP client configuration.
[Agent]
    # - HTTPPort
    # Agent HTTP server port.
    # Default: 19704
    #
    # HTTPPort = 19704

    [Agent.Auth]
        # Must match HTTPServer.Auth.BearerToken on every Node.
        BearerToken = "REPLACE_WITH_RANDOM_HEX"

StoreDSN is the SQLite data source for durable job state. Relative paths are resolved from the configuration directory.

Profiling and tracing use the same quota model but retain independent values. Their resource cost and expected concurrency differ, so a shared limit would allow one workload to starve the other.

Agent request transport safeguards use internal client defaults. Job polling, deadlines, unavailable grace, and retention are owned independently by Jobs.Controller; no runtime negotiation with the Node is performed.

During shutdown, the API server leaves active Node operations running. A replacement API server recovers active Jobs from durable state and resumes monitoring without replaying Start.

5. Elasticsearch/OpenSearch

# Optional Elasticsearch/OpenSearch backend for querying profiling data.
[Elasticsearch]
    # Address, Username, and Password must be configured together to enable
    # this backend.
    #
    # - Address
    # Elasticsearch or OpenSearch HTTP address.
    #
    # - Username
    # Elasticsearch or OpenSearch username.
    #
    # - Password
    # Elasticsearch or OpenSearch password.
    #
    # - Index
    # Index containing huatuo-bamai profiling data.
    # Default: "huatuo_bamai"
    #
    # Address = "https://elasticsearch.example.com:9200"
    # Username = "huatuo-apiserver"
    # Password = "REPLACE_WITH_STRONG_PASSWORD"
    # Index = "huatuo_bamai"

Storage is optional. Address, Username, and Password must either all be empty or all be configured. Index defaults to huatuo_bamai and must match the collector storage index. When disabled, raw-profile and flame graph query routes are not registered.

6. Authentication and Authorization

# Authentication configuration.
[Auth]
    # - ID
    # Stable principal identifier stored with jobs.
    #
    # - BearerToken
    # Secret used only to authenticate requests. IDs and tokens must be unique.
    #
    # - Admin
    # Whether the principal has unrestricted API access.
    #
    # - Permissions
    # API method and path patterns granted to a restricted principal.
    #
    # Administrator example:
    # [[Auth.Users]]
    #     ID = "administrator"
    #     BearerToken = "REPLACE_WITH_RANDOM_HEX"
    #     Admin = true
    #
    # Restricted example:
    # [[Auth.Users]]
    #     ID = "huatuo-front"
    #     BearerToken = "REPLACE_WITH_ANOTHER_RANDOM_HEX"
    #     Permissions = [
    #         "GET /v1/tracing",
    #         "GET /v1/tracing/**",
    #         "GET /v1/profiling",
    #         "GET /v1/profiling/**",
    #     ]
  • ID is the required stable principal identifier stored with jobs.
  • BearerToken is a required secret used only to authenticate requests.
  • Admin grants access to all routes and ignores Permissions.
  • Permissions is required for non-admin users. Entries may be path-only or prefixed with an HTTP method. * matches one path segment and ** matches a suffix.

IDs and bearer tokens must each be unique. Rotating a bearer token does not change job ownership because tokens are never used as principal IDs.

/readyz, /metrics, and /version are public. /debug/pprof/** and /v1/profiling/flamegraph/** require an administrator.

7. Profiling

# Profiling result-link configuration.
[Profiling]
    # - DashboardBaseURL
    # Optional dashboard base URL. Result URLs are omitted when empty.
    # Default: empty
    #
    # DashboardBaseURL = "https://grafana.example.com/d"
  • DashboardBaseURL is optional and must use HTTP or HTTPS when configured. Result-capable jobs omit a dashboard URL when it is empty.

Profiler execution and aggregation settings are Node-local configuration.