# File extraction filters

**URL:** <https://community.zeek.org/t/file-extraction-filters/3201>\
**Category:** Zeek\
**Created:** [July 28, 2014, 9:56pm UTC](https://community.zeek.org/t/file-extraction-filters/3201 "2014-07-28T21:56:19Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mike\_Kolkebeck](https://avatars.discourse-cdn.com/v4/letter/m/a183cd/32.png) [@Mike\_Kolkebeck](https://community.zeek.org/u/Mike_Kolkebeck)\
**Post date:** [July 28, 2014, 9:56pm UTC](https://community.zeek.org/t/file-extraction-filters/3201/1 "2014-07-28T21:56:19Z")

</div>

I have two questions on the file extraction framework:

1) If I only want to capture files from a specific worker or ip ranges, what is the best/simplest way to ensure that this happens?  
-I've tried using f$info$tx\_hosts with event file\_new, but this seems inconsistently populated, and using f$conns with event file\_new seems consistent, but I don't know if it's the best/simplest way.

2) If missing\_bytes \> 0, what is the best/simplest way to remove the file (and possibly clear it from logging a successful extract in the files.log file)?  
-I've tested using event file\_state\_remove, and I can use system to rm the file, but again I'm not sure this is the best/simplest way, and the files.log continues to show this as extracted.

---

<div class="post-metadata">

**Author:** ![Siwek\_Jon](https://avatars.discourse-cdn.com/v4/letter/s/90db22/32.png) [@Siwek\_Jon](https://community.zeek.org/u/Siwek_Jon)\
**Post date:** [July 29, 2014, 2:48pm UTC](https://community.zeek.org/t/file-extraction-filters/3201/2 "2014-07-29T14:48:34Z")

</div>

> I have two questions on the file extraction framework:
> 
> 1) If I only want to capture files from a specific worker or ip ranges, what is the best/simplest way to ensure that this happens?  
> -I've tried using f$info$tx\_hosts with event file\_new, but this seems inconsistently populated, and using f$conns with event file\_new seems consistent, but I don't know if it's the best/simplest way.

In either case, I’d probably try using “file\_over\_new\_connection” instead of “file\_new” — it might end up not mattering for your use, but the fields you’re inspecting are more closely associated with the former event. A given file can technically be transferred over many different connections, depending on the protocol involved, so using “file\_new” may not always give the full story since that’s only ever raised once for a given file.

Using f$info${tx,rx}\_hosts may be better if transfer direction is important, otherwise f$conns should be fine.

> 2) If missing\_bytes \> 0, what is the best/simplest way to remove the file (and possibly clear it from logging a successful extract in the files.log file)?  
> -I've tested using event file\_state\_remove, and I can use system to rm the file, but again I'm not sure this is the best/simplest way, and the files.log continues to show this as extracted.

There’s the “file\_gap” event that you might want to handle, call “Files::remove\_analyzer”, then use a system call to rm the file, and finally “delete f$info$extracted;” to unset the field and prevent it from being logged in files.log.

- Jon

---

<div class="post-metadata">

**Author:** ![Mike\_Kolkebeck](https://avatars.discourse-cdn.com/v4/letter/m/a183cd/32.png) [@Mike\_Kolkebeck](https://community.zeek.org/u/Mike_Kolkebeck)\
**Post date:** [July 29, 2014, 3:38pm UTC](https://community.zeek.org/t/file-extraction-filters/3201/3 "2014-07-29T15:38:05Z")

</div>

Does "file\_over\_new\_connection" fire at the same time as "file\_new" when there is a new file? More specifically, will I ever lose any bytes by using this event over "file\_new"?

---

<div class="post-metadata">

**Author:** ![Siwek\_Jon](https://avatars.discourse-cdn.com/v4/letter/s/90db22/32.png) [@Siwek\_Jon](https://community.zeek.org/u/Siwek_Jon)\
**Post date:** [July 29, 2014, 4:37pm UTC](https://community.zeek.org/t/file-extraction-filters/3201/4 "2014-07-29T16:37:33Z")

</div>

“file\_new” is immediately followed by at least one “file\_over\_new\_connection” (if you’re dealing w/ only files extracted from the network), so there’s not a difference in terms of what bytes have been seen yet. But you may have to think about that event being raised more than once per file and possibly not at the start of a file after the first time, whereas “file\_new” is guaranteed to be once at the start of a file. Not sure which will end up better/simpler for the code you’re writing, but hope that helps explain the differences.

- Jon

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex011/uploads/zeek/original/1X/f09d732bc2cc7c7cc7e35db67cf4e1d5233ce7a7.png) [@system](https://community.zeek.org/u/system)\
**Post date:** [May 6, 2022, 3:41pm UTC](https://community.zeek.org/t/file-extraction-filters/3201/5 "2022-05-06T15:41:56Z")

</div>


