An ingest pipeline runs on the Elasticsearch node before a document is indexed. Use one when the shipper (Filebeat, Logstash, syslog) sends a raw line and you want it parsed without running a separate Logstash filter.
OPNsense syslog pipeline¶
Run this in Kibana Dev Tools:
PUT _ingest/pipeline/opnsense
{
"description": "Parse OPNsense syslog into ECS fields",
"processors": [
{
"grok": {
"field": "message",
"patterns": [
"%{SYSLOGTIMESTAMP:_tmp.timestamp} %{HOSTNAME:observer.hostname} %{DATA:process.name}(?:\\[%{POSINT:process.pid:long}\\])?: %{GREEDYDATA:_tmp.msg}"
]
}
},
{
"date": {
"field": "_tmp.timestamp",
"target_field": "@timestamp",
"formats": ["MMM d HH:mm:ss", "MMM dd HH:mm:ss"],
"timezone": "America/Toronto"
}
},
{
"set": { "field": "observer.type", "value": "firewall" }
},
{
"remove": { "field": "_tmp", "ignore_missing": true }
}
],
"on_failure": [
{
"set": { "field": "error.message", "value": "{{ _ingest.on_failure_message }}" }
}
]
}
What changed compared with a minimal grok + date pipeline:
- Fields go straight to ECS names (see ECS), and there's no
renamestep. timezoneis set. Classic BSD syslog timestamps carry no zone or year, so without it Elasticsearch assumes UTC (and the current year).- Temporary fields live under
_tmpand are removed at the end. on_failurekeeps unparseable lines and records why they failed, instead of rejecting them.
Test before you use it¶
POST _ingest/pipeline/opnsense/_simulate
{
"docs": [
{ "_source": { "message": "Jan 6 18:22:24 fw01 filterlog[12345]: 5,,,1000000103,igb0,match,block,in,4,..." } }
]
}
Check that @timestamp, observer.hostname and process.name are populated and that error.message is absent.
Attach it¶
Either set it as the default pipeline in the index template (index.default_pipeline), which is preferred, or name it per request:
POST logs-opnsense-default/_doc?pipeline=opnsense