<!-- Add the related story/sub-task/bug number, like Resolves #123, or remove if NA --> **Related issue:** Resolves #45715 # Details This PR refactors the way the charts module stores historical data to use the [roaring bitmap](https://github.com/RoaringBitmap/roaring) package instead of saving raw bitmaps. See [this blurb](https://github.com/RoaringBitmap/roaring#how-does-roaring-compares-with-the-alternatives) to learn how roaring compresses data, but TL;DR for our purposes it represents a huge improvement especially for larger deployments where host ID numbers may be very large. In testing, some data was reduced 96%. The majority of the changes in this PR are straight swapping of types from `[]byte` to `*roaring.Bitmap` in vars and function signatures, and updating the internals of our bit math helpers to use roaring methods instead of native AND and OR methods. I've tried to comment on all functional changes. Since the charts have been shipped already, so there will be data in the wild in the prior "dense" format, the code still handles dense bitmaps on _read_, but will always _write_ roaring bitmaps. The majority of the data will therefore have turned over within 30 days on its own, but I plan on a follow-up PR that will transform open rows when the cron runs so that we should be guaranteed to turn over completely within 30 days. # Checklist for submitter If some of the following don't apply, delete the relevant line. - [X] Changes file added for user-visible changes in `changes/`, `orbit/changes/` or `ee/fleetd-chrome/changes`. See [Changes files](https://github.com/fleetdm/fleet/blob/main/docs/Contributing/guides/committing-changes.md#changes-files) for more information. - [X] Input data is properly validated, `SELECT *` is avoided, SQL injection is prevented (using placeholders for values in statements), JS inline code is prevented especially for url redirects, and untrusted data interpolated into shell scripts/commands is validated against shell metacharacters. ## Testing - [X] Added/updated automated tests - Tests updated to accommodate the new format, and existing unchanged tests act as proof against regression - [X] QA'd all new/changed functionality manually - Using a tool that dumps the `host_scd_data` rows data into a JSON file (with the keys being entity_id+data and the values being host IDs on that date), compared the data from main branch and this and confirmed they're identical - With a host count of ~9000, some of which have IDs of over 1,000,000, the data storage requirements were: * 82,558,976 bytes for dense * 2,867,200 for roaring (a 96% decrease) For unreleased bug fixes in a release candidate, one of: - [X] Confirmed that the fix is not expected to adversely impact load test results - should hugely improve - [X] Alerted the release DRI if additional load testing is needed ## Database migrations - [X] Checked schema for all modified table for columns that will auto-update timestamps during migration. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Implemented roaring bitmaps in historical data collection to optimize bitmap handling for chart data aggregation * Added encoding support to bitmap storage schema for flexible data representation <!-- end of auto-generated comment: release notes by coderabbit.ai -->
165 lines
7.2 KiB
Go
165 lines
7.2 KiB
Go
package api
|
|
|
|
import (
|
|
"context"
|
|
"time"
|
|
|
|
"github.com/RoaringBitmap/roaring"
|
|
)
|
|
|
|
// SampleStrategy describes how a dataset's samples combine within a bucket and
|
|
// whether rows can collapse across buckets when the bitmap is unchanged.
|
|
type SampleStrategy string
|
|
|
|
const (
|
|
// SampleStrategyAccumulate means each sample is a partial observation.
|
|
// Writes: every row is born closed (valid_to set at insert time to bucketEnd).
|
|
// Within-bucket samples OR-merge into the existing row via ODKU; a sample in
|
|
// a new bucket just creates a new row with a new valid_from. No explicit
|
|
// close step, no cross-bucket collapse.
|
|
// Reads: bucket value = OR of every row whose interval overlaps the bucket
|
|
// ("hosts observed at any point during the bucket").
|
|
// Used for datasets like uptime and software usage.
|
|
// @todo: implement job to collapse identical consecutive rows
|
|
// to optimize storage and query performance.
|
|
SampleStrategyAccumulate SampleStrategy = "accumulate"
|
|
|
|
// SampleStrategySnapshot means each sample is the full state of a single moment.
|
|
// Writes: rows are always keyed to 1h boundaries (so row transitions align
|
|
// to hour marks regardless of tz). Within a 1h write-bucket, the latest
|
|
// sample's bitmap overwrites via ODKU — last sample wins. Across buckets,
|
|
// unchanged state keeps the row open (valid_to = sentinel); a changed sample
|
|
// closes the prior row at the new hour boundary and opens a new one.
|
|
// Reads: bucket value = OR across entities of each entity's row active at
|
|
// bucketEnd ("state as of the end of the bucket"). An entity whose row was
|
|
// closed mid-bucket with no replacement is absent at bucketEnd.
|
|
// Used for datasets like CVE and software inventory.
|
|
SampleStrategySnapshot SampleStrategy = "snapshot"
|
|
)
|
|
|
|
// Dataset defines the interface for a chartable dataset.
|
|
type Dataset interface {
|
|
// Name returns the dataset identifier used in the DB and API path.
|
|
Name() string
|
|
|
|
// DefaultResolutionHours returns the default display granularity in hours.
|
|
// Used when the caller doesn't specify RequestOpts.Resolution. Unrelated
|
|
// to write-side granularity — all collectors write at 1h regardless of
|
|
// display resolution; see SampleStrategy for details.
|
|
DefaultResolutionHours() int
|
|
|
|
// SampleStrategy returns how samples combine within and across buckets.
|
|
SampleStrategy() SampleStrategy
|
|
|
|
// Collect is called by the cron job to populate data in bulk.
|
|
//
|
|
// disabledFleetIDs scopes which fleets contribute to this collection. The
|
|
// orchestrator derives it from per-team config (teams whose Enabled(name)
|
|
// is false). Implementations should use this to filter out hosts from
|
|
// disabled fleets when collecting data.
|
|
//
|
|
// No-team hosts (team_id IS NULL) are always included when the orchestrator
|
|
// invokes Collect — the orchestrator skips Collect entirely if the global
|
|
// flag is off.
|
|
Collect(ctx context.Context, store DatasetStore, now time.Time, disabledFleetIDs []uint) error
|
|
|
|
// DefaultVisualization returns the default visualization type (e.g. "line", "heatmap").
|
|
DefaultVisualization() string
|
|
}
|
|
|
|
// DatasetStore is the narrow interface that datasets need for their Collect
|
|
// method. It is satisfied by the chart internal Datastore, keeping dataset
|
|
// implementations decoupled from internals.
|
|
type DatasetStore interface {
|
|
// FindOnlineHostIDs returns host IDs that are "online right now" per the
|
|
// product's standard online predicate (host_seen_times.seen_time within
|
|
// the host's own check-in interval). MDM-only mobile devices (iOS,
|
|
// iPadOS, Android) are excluded by design — they don't have
|
|
// host_seen_times rows. Used by datasets like uptime.
|
|
FindOnlineHostIDs(ctx context.Context, now time.Time, disabledFleetIDs []uint) ([]uint, error)
|
|
|
|
// AffectedHostIDsByCVE returns host IDs grouped by CVE, scoped to the given
|
|
// cves set. nil or empty cves returns an empty map — callers must pass the
|
|
// CVE set they want to collect for. Unresolved-only is implicit in the
|
|
// underlying joins: a host's software/OS row transitions when it upgrades
|
|
// past the vulnerable version, so the join naturally stops matching.
|
|
AffectedHostIDsByCVE(ctx context.Context, disabledFleetIDs []uint, cves []string) (map[string][]uint, error)
|
|
|
|
// TrackedCriticalCVEs returns CVE IDs matching the iteration-1 curated
|
|
// filter: critical (CVSS >= 9.0) CVEs on a hard-coded set of software
|
|
// titles, unioned with all critical OS vulnerabilities. Used by the CVE
|
|
// collector to scope collection to only the CVEs the chart actually
|
|
// renders. See TODO in the mysql implementation.
|
|
TrackedCriticalCVEs(ctx context.Context) ([]string, error)
|
|
|
|
// RecordBucketData writes one or more entity bitmaps for the given bucket
|
|
// using the specified sample strategy. See SampleStrategy for semantics.
|
|
// Bitmaps are passed in op form (*roaring.Bitmap); the datastore
|
|
// serializes via chart.BitmapToBlob at the storage boundary.
|
|
RecordBucketData(
|
|
ctx context.Context,
|
|
dataset string,
|
|
bucketStart time.Time,
|
|
bucketSize time.Duration,
|
|
strategy SampleStrategy,
|
|
entityBitmaps map[string]*roaring.Bitmap,
|
|
) error
|
|
}
|
|
|
|
// Host is a minimal host type for authorization checks within the chart bounded context.
|
|
// The JSON tags matter: the OPA rego policy reads object.team_id via the JSON-encoded
|
|
// input, so renaming or dropping the tag silently breaks team-scoped authorization.
|
|
type Host struct {
|
|
ID uint `json:"id"`
|
|
TeamID *uint `json:"team_id"`
|
|
}
|
|
|
|
// AuthzType implements platform_authz.AuthzTyper.
|
|
func (h *Host) AuthzType() string { return "host" }
|
|
|
|
// DataPoint represents a single data point in the chart response.
|
|
type DataPoint struct {
|
|
Timestamp time.Time `json:"timestamp"`
|
|
Value int `json:"value"`
|
|
}
|
|
|
|
// Response is the API response for chart data.
|
|
type Response struct {
|
|
Metric string `json:"metric"`
|
|
Visualization string `json:"visualization"`
|
|
TotalHosts int `json:"total_hosts"`
|
|
Resolution string `json:"resolution"`
|
|
Days int `json:"days"`
|
|
Filters Filters `json:"filters"`
|
|
Data []DataPoint `json:"data"`
|
|
}
|
|
|
|
// RequestOpts captures the parsed query parameters for a chart request.
|
|
type RequestOpts struct {
|
|
Days int
|
|
// Resolution is the display granularity in hours. Must be 0 or a positive
|
|
// divisor of 24. 0 means "use the dataset's default resolution."
|
|
Resolution int
|
|
// TZOffsetMinutes is the client's UTC offset as reported by JavaScript's
|
|
// Date.getTimezoneOffset() (positive = west of UTC, e.g. CDT = 300).
|
|
// Used to align hourly bucket boundaries to local time.
|
|
TZOffsetMinutes int
|
|
// TeamID scopes the request to a single team. nil = global (authz + data
|
|
// both fall back to the user's accessible scope). *TeamID == 0 means
|
|
// hosts with no team assignment, matching Fleet's convention elsewhere.
|
|
TeamID *uint
|
|
LabelIDs []uint
|
|
Platforms []string
|
|
IncludeHostIDs []uint
|
|
ExcludeHostIDs []uint
|
|
}
|
|
|
|
// Filters captures the applied filters for a chart request.
|
|
type Filters struct {
|
|
TeamID *uint `json:"fleet_id,omitempty"`
|
|
LabelIDs []uint `json:"label_ids,omitempty"`
|
|
Platforms []string `json:"platforms,omitempty"`
|
|
IncludeHostIDs []uint `json:"include_host_ids,omitempty"`
|
|
ExcludeHostIDs []uint `json:"exclude_host_ids,omitempty"`
|
|
}
|