# Inaccuracy of orig\_bytes and resp\_bytes?

**URL:** https://community.zeek.org/t/inaccuracy-of-orig-bytes-and-resp-bytes/6804
**Category:** Zeek
**Created:** [November 19, 2022, 3:33pm UTC](https://community.zeek.org/t/inaccuracy-of-orig-bytes-and-resp-bytes/6804 "2022-11-19T15:33:17Z")
**Posts on this page:** 4
**Page:** 1

<div class="post-metadata">

### Author: ![ryanwmcronald](https://avatars.discourse-cdn.com/v4/letter/r/50afbb/32.png) [@ryanwmcronald](https://community.zeek.org/u/ryanwmcronald)
#### Post date: [November 19, 2022, 3:33pm UTC](https://community.zeek.org/t/inaccuracy-of-orig-bytes-and-resp-bytes/6804/1 "2022-11-19T15:33:17Z")

</div>

WIth reference to the conn.log spec [base/protocols/conn/main.zeek — Book of Zeek (git/master)](https://docs.zeek.org/en/master/scripts/base/protocols/conn/main.zeek.html)

_ **orig\_bytes** _: _The number of payload bytes the originator sent. For TCP this is taken from sequence numbers and **might be inaccurate** (e.g., due to large connections)  
 **resp\_bytes** : The number of payload bytes the responder sent. See orig\_bytes_

Are these values inaccurate to the point that they should not be relied upon? Here is an example of a common conn.log entry for TCP traffic:

```auto
orig_bytes: 2,145,967,013
resp_bytes: 81,091,225
duration: 15.638
orig_ip_bytes: 104
resp_ip_bytes: 340

```

The `orig_bytes` and `resp_bytes` values appear ‘impossible’ due to the amount of bytes, the short duration, and in relation to the respective ip\_bytes. Such entries are really messing with visualizations and analysis using `orig_bytes` and `resp_bytes`. Are others relying on `orig_ip_bytes` and `resp_ip_bytes` instead of `orig_bytes` and `resp_bytes`? Should `orig_bytes` ever be larger than `orig_ip_bytes`?

We are running Zeek 5.0.3 on a RPM/RedHat-based multi-node cluster with pf\_ring, Myricom 10Gb NICs, OpenSearch, and Corelight’s [zeek2es](https://github.com/corelight/zeek2es).

Any comments on the experiences of others with these records would be appreciated.

Ryan

---

<div class="post-metadata">

### Author: ![ryanwmcronald](https://avatars.discourse-cdn.com/v4/letter/r/50afbb/32.png) [@ryanwmcronald](https://community.zeek.org/u/ryanwmcronald)
#### Post date: [November 21, 2022, 3:37am UTC](https://community.zeek.org/t/inaccuracy-of-orig-bytes-and-resp-bytes/6804/2 "2022-11-21T03:37:53Z")

</div>

More on this topic,… on a sample of 37.4 million `"proto": "tcp"` conn.log entries approximately 30% have `orig_bytes` larger than `orig_ip_bytes`.

Considering the conn.log spec:

_ **orig\_ip\_bytes** _: _Number of IP level bytes that the originator sent (as seen on the wire, taken from the [IP total\_length](https://en.wikipedia.org/wiki/IPv4#Total_Length) header field)._

this should always be larger than `"orig_bytes"` by a minimum of `"orig_pkts"` \* 20 bytes for the minimum size of an IPv4 header.

Here are more details on the example given above:

```auto
    "proto": "tcp",
    "duration": 15.637575,
    "orig_bytes": 2145967013,
    "resp_bytes": 81091225,
    "conn_state": "RSTR",
    "local_orig": false,
    "local_resp": true,
    "missed_bytes": 81091225,
    "history": "ShAgr",
    "orig_pkts": 2,
    "orig_ip_bytes": 104,
    "resp_pkts": 6,
    "resp_ip_bytes": 340,

```

If `"orig_pkts": 2` is to be trusted, the theoretical maximum for `"orig_bytes"` is 65,535 minus two 20 byte IP headers being 65,495 bytes. Here it is showing 2,145,967,013 bytes.

If `"resp_pkts": 6` is to be trusted, the theoretical maximim for `"resp_bytes"` is 65,535 minus six 20 byte IP headers being 65,475 bytes. Here it is showing 81,091,225 bytes.

I notice most of the impacted conn.log entries have `"conn_state": "RSTR"` and `"history"` contains a `"g"` (for gaps) and a large value for `"missed_bytes"`.

Perhaps there should be a condition in the code that if the value computed for `"orig_bytes"` or `"resp_bytes"` is larger than the number of IPv4 packets \* 65,335 bytes then to set the corresponding \_bytes value to “-1” (unknown) instead of logging impossibly large values.

Ryan

---

<div class="post-metadata">

### Author: ![awelzel](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/awelzel/32/609_2.png) [@awelzel](https://community.zeek.org/u/awelzel)
#### Post date: [November 23, 2022, 12:05pm UTC](https://community.zeek.org/t/inaccuracy-of-orig-bytes-and-resp-bytes/6804/3 "2022-11-23T12:05:24Z")

</div>

@ryanwmcronald - thanks for the data and observations.

> Should `orig_bytes` ever be larger than `orig_ip_bytes`?

The `orig_bytes` and `resp_bytes` for TCP are determined from the TCP sequence/ack numbers within packets, while the `orig_ip_bytes` and `resp_ip_bytes` accounts for packets actually seen by Zeek on the wire, as the docs state. So yes, in the presence of packet loss/gaps (as your entry shows) that can happen.

> More on this topic,… on a sample of 37.4 million `"proto": "tcp"` conn.log entries approximately 30% have `orig_bytes` larger than `orig_ip_bytes`.

Do you incur packet loss? Can try the following [stats](https://docs.zeek.org/en/master/scripts/policy/misc/stats.zeek.html) and [capture-loss](https://docs.zeek.org/en/master/scripts/policy/misc/capture-loss.zeek.html) policy scripts or calculate the percentage of conn entries with gaps in them yourself.

You could further look into weird.log for the affected connections and see if something stands out.

If it appears unrealistic that a 2GB TCP connection between the two involved hosts above should happen, would it be possible to capture and share an anonymized pcap containing the few packets of such a connection that Zeek actually saw (I know this isn’t easy). That would allow to dig into the details and possibly explain what’s going on.

---

<div class="post-metadata">

### Author: ![system](https://canada1.discourse-cdn.com/flex011/uploads/zeek/original/1X/f09d732bc2cc7c7cc7e35db67cf4e1d5233ce7a7.png) [@system](https://community.zeek.org/u/system)
#### Post date: [April 8, 2026, 7:49pm UTC](https://community.zeek.org/t/inaccuracy-of-orig-bytes-and-resp-bytes/6804/4 "2026-04-08T19:49:33Z")

</div>

This topic was automatically closed 2 days after the last reply. New replies are no longer allowed.
