Add support for publishing logs to NATS. - #36527
Conversation
…sult in published messages.
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #36527 +/- ##
==========================================
- Coverage 65.86% 65.85% -0.02%
==========================================
Files 2361 2363 +2
Lines 187305 187608 +303
Branches 8010 7976 -34
==========================================
+ Hits 123364 123540 +176
- Misses 52664 52766 +102
- Partials 11277 11302 +25
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Sentry. 🚀 New features to boost your workflow:
|
|
@ebusto thanks for your contribution! Temporarily converting this to a draft PR while we bring this change through our Drafting process. You'll be notified on the connected issue once it makes it into a release. |
There was a problem hiding this comment.
LGTM ✅ . I have a few comments in server/logging/nats.go, but nothing major.
I've tested this locally as follows (will document it in a follow-up PR):
- Installed
nats:go install github.com/nats-io/natscli/nats@latest - Started the NATS server locally:
nats-server:
- Start fleet with the following flags:
FLEET_ACTIVITY_ENABLE_AUDIT_LOG=true FLEET_ACTIVITY_AUDIT_LOG_PLUGIN=nats FLEET_OSQUERY_RESULT_LOG_PLUGIN=nats FLEET_OSQUERY_STATUS_LOG_PLUGIN=nats FLEET_NATS_SERVER=nats://localhost:4222 FLEET_NATS_STATUS_SUBJECT=osquery_status FLEET_NATS_RESULT_SUBJECT=osquery_result FLEET_NATS_AUDIT_SUBJECT=fleet_audit ./build/fleet serve --dev- Configured a scheduled query:
- On a separate terminal, ran
./nats --server=nats://localhost:4222 subscribe ">". Per GPT, thislistens on all subjects and prints any messages that get published. Eventually, I saw the query results being logged:
EDIT: I also tested with one of the authentication methods, NKey:
- Install
nkey:go install github.com/nats-io/nkeys/nk@latest. - Run
nk -gen user -pubout. This will output something like:
SUxxx
Uyyy
- Copy the output above to a txt file, e.g.
nkey-cred-file.txt. - Create a new NATS server config file, such as
nats-server-config.confwith this content:
authorization {
users = [
{
nkey: "Uxxx"
}
]
}
- Start the NATS server providing the config file above:
./nats-server -config nats-server-config.conf. You should see a log sayingUsing configuration file: nats-server-config.conf.
- Start fleet specifying the NKey cred file to use:
FLEET_ACTIVITY_ENABLE_AUDIT_LOG=true FLEET_ACTIVITY_AUDIT_LOG_PLUGIN=nats FLEET_OSQUERY_RESULT_LOG_PLUGIN=nats FLEET_OSQUERY_STATUS_LOG_PLUGIN=nats FLEET_NATS_SERVER="nats://localhost:4222" FLEET_NATS_STATUS_SUBJECT=osquery_status FLEET_NATS_RESULT_SUBJECT=osquery_result FLEET_NATS_AUDIT_SUBJECT=fleet_audit FLEET_NATS_NKEY_FILE="nkey-cred-file.txt" ./build/fleet serve --dev
| // Define the supported compression algorithms. | ||
| var compressionOk = map[string]bool{ | ||
| "gzip": true, | ||
| "snappy": true, | ||
| "zstd": true, | ||
| } |
There was a problem hiding this comment.
I'd remove the comment on L66 and rename the variable to a more intent-revealing name, e.g. supportedCompressionAlgorithms.
Also, this could probably be a slice instead of a map[string]bool. Since all entries are effectively true, the boolean doesn’t add much value, and a list of supported algorithms feels more readable and intuitive here. With so few elements, performance isn’t really a concern IMO.
There was a problem hiding this comment.
Idiomatic Go for sets is map[string]struct{}.
| "zstd": true, | ||
| } | ||
|
|
||
| // NewNatsLogWriter creates a new NATS log writer. |
There was a problem hiding this comment.
The method's name already makes me infer that it does this, so I'd remove this comment.
| // Whether to use JetStream. | ||
| jetstream bool |
There was a problem hiding this comment.
since this is a boolean I'd prefix it with is or has, or maybe in this case useJetstream
| @@ -0,0 +1,484 @@ | |||
| package logging | |||
There was a problem hiding this comment.
IMO, some of the comments in this file shouldn't be needed since the reader can infer what the code does by reading the variable names or following through the statements. For example, the comment above NewNatsLogWriter says // NewNatsLogWriter creates a new NATS log writer., which I don't think adds any value and is also redundant.
iansltx
left a comment
There was a problem hiding this comment.
Approving FE for LogDestinationIndicator.tsx and frontend/interfaces/config.ts.
|
@ebusto this change shipped in Fleet 4.80.0. Thanks for adding this feature! 🎉 |
**Related issue:** Resolves #25574 # Checklist for submitter - [x] Changes file added for user-visible changes in `changes/` - [x] Input data is properly validated, `SELECT *` is avoided, SQL injection is prevented (using placeholders for values in statements), JS inline code is prevented especially for url redirects, and untrusted data interpolated into shell scripts/commands is validated against shell metacharacters. - [x] Timeouts are implemented and retries are limited to avoid infinite loops ## Testing - [x] Added/updated automated tests - [x] QA'd all new/changed functionality manually --- ## Summary - Adds a new `splunk` log plugin that sends osquery logs directly to Splunk's HTTP Event Collector (HEC) endpoint - Eliminates the need for middleware like AWS Firehose when using Splunk as a log destination - Follows the same pattern as existing log destinations (Firehose, Kafka REST, NATS, etc.) - Includes `insecure_skip_verify` option for environments with self-signed TLS certs ## UI changes Follows the same pattern as the NATS log destination PR (#36527) -- adding "Splunk" to the display name, tooltip, and TypeScript type union. No new components, pages, or styles. ### Manage automations modal -- "Log destination: Splunk" <img width="822" height="527" alt="image" src="https://github.com/user-attachments/assets/2533207f-fa95-4364-8ee0-3c39cd3e8e4d" /> ### Query details page -- "Log destination: Splunk" <img width="1905" height="662" alt="image" src="https://github.com/user-attachments/assets/069a5005-f95c-4562-a819-fd8bdcc349f7" /> ### Tooltip on hover <img width="639" height="348" alt="image" src="https://github.com/user-attachments/assets/809a47a6-b82a-4f45-b731-77b2d2c87947" /> ### Edit query form -- "sent to your log destination: Splunk" <img width="451" height="814" alt="image" src="https://github.com/user-attachments/assets/b78b9a57-1f0c-4413-8b7c-654de1fd40a2" /> ### Save new query modal -- "sent to your log destination: Splunk" <img width="536" height="698" alt="image" src="https://github.com/user-attachments/assets/d0a0ab01-66fe-4d63-9190-9c5e840e456d" /> --- ### How it works The Splunk writer (`server/logging/splunk.go`) implements the `fleet.JSONLogger` interface. On startup it performs a health check against the HEC `/services/collector/health` endpoint. On each `Write()` call, it wraps each log entry in Splunk's HEC event format (adding `time`, `index`, `source`, `sourcetype`), batches them up to 1 MB, and POSTs to `/services/collector/event` with the `Authorization: Splunk <token>` header. If a batch exceeds 1 MB it flushes and starts a new one. Events over 1 MB are dropped with a log warning. Transient errors (HTTP 503) are retried with exponential backoff (up to 8 retries). ### Configuration ```yaml osquery: status_log_plugin: splunk result_log_plugin: splunk splunk: url: https://splunk.example.com:8088 token: <HEC token> index: main source: fleet source_type: fleet:json insecure_skip_verify: false # set true for self-signed certs ``` Or via environment variables: ``` FLEET_OSQUERY_STATUS_LOG_PLUGIN=splunk FLEET_OSQUERY_RESULT_LOG_PLUGIN=splunk FLEET_SPLUNK_URL=https://splunk.example.com:8088 FLEET_SPLUNK_TOKEN=<HEC token> FLEET_SPLUNK_INDEX=main FLEET_SPLUNK_SOURCE=fleet FLEET_SPLUNK_SOURCE_TYPE=fleet:json ``` ### Files changed - `server/logging/splunk.go` -- Splunk HEC log writer with batching, retry, and health check - `server/logging/splunk_test.go` -- 9 unit tests - `server/logging/splunk_integration_test.go` -- 3 integration tests against real Splunk (gated by env var) - `server/logging/logging.go` -- Added `SplunkConfig` and `case "splunk"` to factory - `server/config/config.go` -- Added `SplunkConfig` struct and config flags - `cmd/fleet/logging.go` -- Wired Splunk config into logging builder - `server/fleet/app.go` -- Added `SplunkConfig` type for API responses (excludes token) - `server/service/service_appconfig.go` -- Added `case "splunk"` to logging plugin validation - `frontend/interfaces/config.ts` -- Added `"splunk"` to LogDestination type - `frontend/components/LogDestinationIndicator/LogDestinationIndicator.tsx` -- Added Splunk display name and tooltip - `docs/Configuration/fleet-server-configuration.md` -- Splunk config documentation - `docs/Get started/FAQ.md` -- Updated plugin list - `articles/log-destinations.md` -- Updated Splunk section with native HEC docs - `changes/25574-splunk-log-destination` -- Change file ## Test plan ### Unit tests (9 tests) - [x] `TestSplunkWrite` -- sends 3 events, verifies HEC format, auth header, index/source/sourcetype - [x] `TestSplunkWriteEmpty` -- empty logs don't trigger HTTP request - [x] `TestSplunkServerError` -- HEC 403 propagates as error - [x] `TestSplunkHealthCheckFailure` -- constructor fails on bad health - [x] `TestSplunkRecordTooBig` -- oversized events (>1MB) are dropped, normal events still sent - [x] `TestSplunkSplitBatchBySize` -- logs exceeding 1MB batch limit are split into multiple requests - [x] `TestSplunkRetryOnServiceUnavailable` -- 503 retried with backoff, succeeds on 3rd attempt - [x] `TestSplunkRetryExhausted` -- after 9 attempts (1 + 8 retries) returns error - [x] `TestSplunkMissingConfig` -- empty URL/token returns descriptive error ### Integration tests (3 tests, gated by `SPLUNK_INTEGRATION_TEST=1`) - [x] `TestSplunkIntegration` -- 3 events sent via writer, queried back from Splunk REST API - [x] `TestSplunkIntegrationBatch` -- 100 events in one Write(), all confirmed indexed - [x] `TestSplunkIntegrationBadToken` -- bad token Write() returns 403 ### End-to-end test (macOS ARM64, real osquery agent) 1. Started Splunk Enterprise, MySQL, Redis via Docker 2. Built Fleet server from this branch with `--osquery_status_log_plugin=splunk` 3. Set up Fleet, enrolled a real osquery 5.23.0 agent on this MacBook 4. **83 real osquery status log events indexed in Splunk** with correct source/sourcetype/index 5. Each event contained full osquery data (`hostIdentifier`, `host_uuid`, `calendarTime`, `severity`, `message`, `decorations`) ### Splunk showing real osquery events from Fleet <img width="1910" height="861" alt="image" src="https://github.com/user-attachments/assets/192490bf-d594-4424-a3e3-a18306892873" /> Generated with [Claude Code](https://claude.ai/code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added native Splunk HEC logging destination for status, result, and audit logs. * Updated the log destination UI to display **Splunk** with a dedicated tooltip. * Added Splunk HEC configuration (URL/token/index/source/source type) including TLS verification control. * **Bug Fixes** * Improved log delivery with batching, retries for temporary HTTP failures, and safeguards for oversized events. * **Tests** * Added unit tests and optional integration tests covering routing, batching, retries, and error scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Related issue: Resolves 34890
Checklist for submitter
changes/,orbit/changes/oree/fleetd-chrome/changes.Testing
New Fleet configuration settings
Looking at other log destinations, I couldn't find anything relevant in GitOps. Please let me know if I missed something, however.
fleetd/orbit/Fleet Desktop
I've tested this on both Linux and MacOS.