# \[Proposal\] Language extensions for better Broker support

**URL:** <https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526>\
**Category:** Development\
**Tags:** development\
**Created:** [December 2, 2016, 3:39am UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526 "2016-12-02T03:39:29Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![robin](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/robin/32/599_2.png) [@robin](https://community.zeek.org/u/robin)\
**Post date:** [December 2, 2016, 3:39am UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/1 "2016-12-02T03:39:29Z")

</div>

Bro's current Broker framework has a few pretty inelegant API parts  
because Bro's scripting language doesn't support some of its  
operations well currently. I've put some thoughts together on  
potential language extensions to improve the situation and come to a  
nicer Broker framework API:

&nbsp;&nbsp;&nbsp;&nbsp;[https://www.bro.org/development/projects/broker-lang-ext.html](https://www.bro.org/development/projects/broker-lang-ext.html)

Feedback welcome, this is just a first draft.

Robin

---

<div class="post-metadata">

**Author:** ![Azoff\_Justin\_S](https://avatars.discourse-cdn.com/v4/letter/a/dec6dc/32.png) [@Azoff\_Justin\_S](https://community.zeek.org/u/Azoff_Justin_S)\
**Post date:** [December 2, 2016, 2:19pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/2 "2016-12-02T14:19:12Z")

</div>

Asynchronous executions without when: yes!

Was just talking to Vlad about this yesterday. The examples get even worse as soon as you need to do more than one broker operation in sequence. Something like this with when statements would be unmaintainable:

#Check to see if any of these keys exists  
local v1 = Broker::lookup(h, 41);  
local v2 = Broker::lookup(h, 42);  
local v3 = Broker::lookup(h, 43);

if (v1 || v2 || v3) {  
&nbsp;&nbsp;&nbsp;&nbsp;Broker::set(h, ...)  
}

Or I could see a trivial example like this for counting things per day:

event connection\_established(c: connection)  
{  
&nbsp;&nbsp;# not hardcoded ideally..  
&nbsp;&nbsp;Broker::inc(h, fmt("connections:2016-12-02:addr:%s, c$id$orig\_h), 1);  
&nbsp;&nbsp;Broker::inc(h, fmt("connections:2016-12-02:addr:%s, c$id$resp\_h), 1);  
&nbsp;&nbsp;Broker::inc(h, fmt("connections:2016-12-02:port:%s, c$id$orig\_p), 1);  
&nbsp;&nbsp;Broker::inc(h, fmt("connections:2016-12-02:port:%s, c$id$resp\_p), 1);  
}

Having a way to send a batch of operations would be nice, but that's a separate issue 🙂

---

<div class="post-metadata">

**Author:** ![robin](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/robin/32/599_2.png) [@robin](https://community.zeek.org/u/robin)\
**Post date:** [December 2, 2016, 3:52pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/3 "2016-12-02T15:52:27Z")

</div>

> Something like this with when statements would be  
> unmaintainable:

Yeah, exactly.

> local v1 = Broker::lookup(h, 41);  
> local v2 = Broker::lookup(h, 42);  
> local v3 = Broker::lookup(h, 43);

> Having a way to send a batch of operations would be nice, but that's a separate issue 🙂

Good thought, those three lookups above could all proceed in parallel,  
and then one would wait for them all to finish before continuing.

Robin

---

<div class="post-metadata">

**Author:** ![Siwek\_Jon](https://avatars.discourse-cdn.com/v4/letter/s/90db22/32.png) [@Siwek\_Jon](https://community.zeek.org/u/Siwek_Jon)\
**Post date:** [December 2, 2016, 5:00pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/4 "2016-12-02T17:00:00Z")

</div>

Looks like a big improvement, some ideas:

Not that syntax is super important to nail down right away, but to me, “v as T” syntax is simpler than "as\<T\>(v)”.

Is there a reason "type(v)” can’t be stored? I’d probably find it more intuitive if it could, else I can see myself forgetting or making mistakes related to that. Alternative to even providing “type(v)”, you could have a “v is T” operation and to use an explicit/new “typeswitch (v)” statement instead of re-using “switch (type(v))”.

Between the two switch syntaxes you gave, maybe provide both and allow mixing of syntax between cases. E.g.:

switch ( type(v) ) {  
&nbsp;&nbsp;&nbsp;&nbsp;case bool -\> b:  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;print “it’s a bool, and I need to inspect the value", b;  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;break;

&nbsp;&nbsp;&nbsp;&nbsp;case count:  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;print “it’s a count, but I don’t care what the value is, I just wanted to know if it’s a count”;  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;break;  
}

"Asynchronous executions without when”: I’d go with an explicit keyword to denote when a handler will yield execution — it makes it easier to understand the behavior of a function at a glance without requiring a perfect memory for what functions cause execution to yield. Full co-routine support is a nice step to take and this is already similar enough that adding a keyword like “yield” at this point may make it easier to evolve/plan the language into having generalized co-routine support. It might even be useful to try and spec out the co-routine support now and see how this specific use-case fits in with that.

- Jon

---

<div class="post-metadata">

**Author:** ![Matthias\_Vallentin1](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/matthias_vallentin1/32/596_2.png) [@Matthias\_Vallentin1](https://community.zeek.org/u/Matthias_Vallentin1)\
**Post date:** [December 2, 2016, 5:17pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/5 "2016-12-02T17:17:22Z")

</div>

> Feedback welcome, this is just a first draft.

I like this. Some initial feedback:

&nbsp;&nbsp;- In the switch statement, I would require that a user either provide  
&nbsp;&nbsp;&nbsp;&nbsp;a default case or fully enumerates all cases. Otherwise it's too  
&nbsp;&nbsp;&nbsp;&nbsp;easy to cause harder-to-debug run-time errors.

&nbsp;&nbsp;- For the abbreviated switch statement, what do you think about this  
&nbsp;&nbsp;&nbsp;&nbsp;variation?

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;switch ( type(v) ) {  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;case b: bool  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;print "bool", b;  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;break;

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;case s: set[int]  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;print "set", s;  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;break;

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;case r: some\_record  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;print "record", r;  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;break;

&nbsp;&nbsp;&nbsp;&nbsp;Since we don't use the arrow operator anywhere and already declare  
&nbsp;&nbsp;&nbsp;&nbsp;variable type with colons as in Pascal, this could feel more natural  
&nbsp;&nbsp;&nbsp;&nbsp;to Bro.

&nbsp;&nbsp;- Why do we need a string comparison here?  
&nbsp;&nbsp;&nbsp;&nbsp;  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;if ( type(v) == string )

&nbsp;&nbsp;&nbsp;&nbsp;Wouldn't it suffice to have

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;if ( type(v) == X )

&nbsp;&nbsp;&nbsp;&nbsp;where X is a type instance (e.g., addr, count, set[int]) or a  
&nbsp;&nbsp;&nbsp;&nbsp;user-defined type (e.g., connection)? In other words, I don't  
&nbsp;&nbsp;&nbsp;&nbsp;understand why equality comparison need to be between a type and a  
&nbsp;&nbsp;&nbsp;&nbsp;string.

- Uplevelling, why do we need type() in the first place? Can't we just  
&nbsp;&nbsp;&nbsp;overload the switch statement to work on types as well? Or do you  
&nbsp;&nbsp;&nbsp;have other use cases in mind?

- Regarding asynchronous execution:

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;local h = async lookup\_hostname(“www.icir.org”);

&nbsp;&nbsp;&nbsp;I like the explicit keyword here to signal asynchrony. Ignoring the  
&nbsp;&nbsp;&nbsp;when statement for now, users rely on non-preempted execution both  
&nbsp;&nbsp;&nbsp;within a function and an event handler, and I would argue that many  
&nbsp;&nbsp;&nbsp;scripts are built around this assumption. When these semantics  
&nbsp;&nbsp;&nbsp;change, it could become harder to reason about the execution model.  
&nbsp;&nbsp;&nbsp;However, if we make timeouts mandatory, I woudn't mind dropping the  
&nbsp;&nbsp;&nbsp;async keyword.

&nbsp;&nbsp;&nbsp;&nbsp;Matthias

---

<div class="post-metadata">

**Author:** ![Jan](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@Jan](https://community.zeek.org/u/Jan)\
**Post date:** [December 2, 2016, 8:18pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/6 "2016-12-02T20:18:49Z")

</div>

I really like the suggested improvements and I agree to Jon and Matthias  
regarding the use of "async" to make corresponding calls explicit. The  
first thing I thought of was the get\_current\_packet() or the  
get\_current\_packet\_header() function. If there is some asynchronous  
execution in context of a "new" packet, a transparent syntax might  
obfuscate what is going on there. I also like the switch syntax Matthias  
suggested.

Jan

---

<div class="post-metadata">

**Author:** ![robin](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/robin/32/599_2.png) [@robin](https://community.zeek.org/u/robin)\
**Post date:** [December 4, 2016, 8:40pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/7 "2016-12-04T20:40:06Z")

</div>

Thanks for the feedback. I'm replying below to all three of you, as  
it's all related. I'm hearing strong support for making asynchronous  
calls explicit, which is fine with me.

> “v as T” syntax is simpler than "as\<T\>(v)”.

I see one disadvantage with the "as" syntax: the whitespace will lead  
to additional parentheses being required in some contexts. For  
example, with T being a record: "(v as T)$foobar" vs  
"as\<T\>(v)$foobar". Not necessarily a showstopper though. The "as"  
syntax does feel more in line with other Bro syntax.

Are there other ideas to express casts?

> Is there a reason "type(v)” can’t be stored? I’d probably find it  
> more intuitive if it could, else I can see myself forgetting or making  
> mistakes related to that.

Mind elaborating how you would expect to use that? My reasoning for  
foregoing storage was that I couldn't really see much use for having a  
type stored in variable in the first place because there aren't really  
any operations that you could then use on that variable (other than  
generic ones, like printing it).

> &nbsp;&nbsp;Alternative to even providing “type(v)”, you could have a “v is T”

I like that. That avoids the question of storing types altogether.

> &nbsp;&nbsp;operation and to use an explicit/new “typeswitch (v)” statement  
> &nbsp;&nbsp;instead of re-using “switch (type(v))”.

Hmm ... I'd prefer not introduce a new switch statement. However,  
seems the existing switch could just infer that it's maching types by  
looking at case values: if they are types, it'd use "is" for  
comparision.

> &nbsp;&nbsp;&nbsp;&nbsp;case bool -\> b:  
> &nbsp;&nbsp;&nbsp;&nbsp;case count:

Good point, I like that.

> It might even be useful to try and spec out the co-routine support now  
> and see how this specific use-case fits in with that.

Need to think about that, would clearly be nice if we could pave the  
way for that already.

> &nbsp;&nbsp;- In the switch statement, I would require that a user either provide  
> &nbsp;&nbsp;&nbsp;&nbsp;a default case or fully enumerates all cases. Otherwise it's too  
> &nbsp;&nbsp;&nbsp;&nbsp;easy to cause harder-to-debug run-time errors.

Full enumeration doesn't seem possible, that would mean all possible  
types, no?

Are we requiring "default" for the standard switch statement? If so, I  
agree it makes sense to do the same here (otherwise, not so sure,  
because of consistency).

> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;case b: bool  
> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;case s: set[int]  
> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;case r: some\_record

Interesting thought. I had this at first:

&nbsp;&nbsp;&nbsp;&nbsp;case b: bool: ...  
&nbsp;&nbsp;&nbsp;&nbsp;case s: set[int]: ...  
&nbsp;&nbsp;&nbsp;&nbsp;case r: some\_record: ...

That mimics the existing syntax more closely, but ends up being ugly  
with the two colons. I kind of like your idea, even though it makes  
the syntax a bit non-standard. However, I'm not sure the parser could  
deal with that easily (because there's no clear separation between the  
type and the following code). Also, this might not fit so well with  
Jon's idea of making the identifier optional.

> &nbsp;&nbsp;- Why do we need a string comparison here?  
> &nbsp;&nbsp;&nbsp;&nbsp;  
> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;if ( type(v) == string )

I believe you misread this: "string" \*is\* a type.

> - Uplevelling, why do we need type() in the first place? Can't we just  
> &nbsp;&nbsp;&nbsp;overload the switch statement to work on types as well? Or do you  
> &nbsp;&nbsp;&nbsp;have other use cases in mind?

My main other use case was offering if-style comparision as well for  
types (see previous point). But if we do "is" for that, we indeed  
wouldn't need the type() anymore.

> &nbsp;&nbsp;&nbsp;However, if we make timeouts mandatory, I woudn't mind dropping the  
> &nbsp;&nbsp;&nbsp;async keyword.

I'm going back and forth on whether to make timeouts mandatory. I  
think we need to have a default timeout in any case, we cannot have  
calls linger around forever. But explicitly specifying one each time  
takes away some of the simplicity. On the other hand, with a default  
timeout we'd probably need to throw runtime errors instead of  
returning something, and Bro isn't good with tons of runtime errors  
(which could happen here).

---

<div class="post-metadata">

**Author:** ![Jan](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@Jan](https://community.zeek.org/u/Jan)\
**Post date:** [December 4, 2016, 11:50pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/8 "2016-12-04T23:50:05Z")

</div>

> > The first thing I thought of was the get\_current\_packet() or the  
> > get\_current\_packet\_header() function.
> 
> Can you elaborate? These functions aren't asynchronous currently.  
> Would you change them to being so; and if so, what would that do?

That was just another example for a situation in which "hidden"  
asynchronous execution could easily lead to unintended behavior. Calls  
of get\_current\_packet() in context of the same event handler will return  
different results in case the execution gets interrupted due to an async  
call. But that's most likely an esoteric edge case. Sorry for the noise 🙂

Jan

---

<div class="post-metadata">

**Author:** ![robin](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/robin/32/599_2.png) [@robin](https://community.zeek.org/u/robin)\
**Post date:** [December 5, 2016, 3:51pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/9 "2016-12-05T15:51:18Z")

</div>

Ah, got it, I had misread what you meant. Thanks,

Robin

---

<div class="post-metadata">

**Author:** ![Matthias\_Vallentin1](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/matthias_vallentin1/32/596_2.png) [@Matthias\_Vallentin1](https://community.zeek.org/u/Matthias_Vallentin1)\
**Post date:** [December 6, 2016, 11:17am UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/10 "2016-12-06T11:17:05Z")

</div>

> \> Alternative to even providing “type(v)”, you could have a “v is T”
> 
> I like that. That avoids the question of storing types altogether.

I think we can cover all type-based dispatching with "is" and "switch."  
In fact, I see "x is T" as syntactic sugar for:

function is(x: any) {  
&nbsp;&nbsp;&nbsp;&nbsp;switch(x) {  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;default:  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return false;  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;case T:  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return true;  
&nbsp;&nbsp;&nbsp;&nbsp;}  
}

> \> It might even be useful to try and spec out the co-routine support now  
> \> and see how this specific use-case fits in with that.
> 
> Need to think about that, would clearly be nice if we could pave the  
> way for that already.

+1

> Full enumeration doesn't seem possible, that would mean all possible  
> types, no?

Yeah, not in a distributed systems, at least. I think I was coming from  
C++ here where dispatching on type-safe unions requires complete  
enumeration of all overloads. In Bro, we don't need that because we  
cannot know the complete list of types a priori.

> Are we requiring "default" for the standard switch statement? If so, I  
> agree it makes sense to do the same here (otherwise, not so sure,  
> because of consistency).

Sounds good to me.

> \> if ( type(v) == string )
> 
> I believe you misread this: "string" \*is\* a type.

Ah, now that makes sense!

> \> - Uplevelling, why do we need type() in the first place? Can't we just  
> \> overload the switch statement to work on types as well? Or do you  
> \> have other use cases in mind?
> 
> My main other use case was offering if-style comparision as well for  
> types (see previous point). But if we do "is" for that, we indeed  
> wouldn't need the type() anymore.

Good, I like "is" better than a type function because of readability.

> But explicitly specifying one each time takes away some of the  
> simplicity.
> 
> On the other hand, with a default timeout we'd probably need to throw  
> runtime errors instead of returning something, and Bro isn't good with  
> tons of runtime errors (which could happen here).

How is a timeout different from a failed computation due to a different  
reason (e.g., connection failure, store backend error)? I think we need  
to consider errors as possible outcomes of \*any\* asynchronous operation.  
More generally, we need well-defined semantics for composing  
asynchronous computations. For example, what do we do here?

&nbsp;&nbsp;&nbsp;&nbsp;# f :: T -\> U  
&nbsp;&nbsp;&nbsp;&nbsp;# g :: T  
&nbsp;&nbsp;&nbsp;&nbsp;local x = async f(async g())?

One solution: if g fails, x should contain g's error---while f will  
never execute. But in the common case of success, the user just wants to  
work with x as an instance of T, e.g., when T = count:

&nbsp;&nbsp;&nbsp;&nbsp;count y = 42;  
&nbsp;&nbsp;&nbsp;&nbsp;print x + y; # use x as type count

If x holds an error, then Bro would raise a runtime error in the  
addition operator. Would that make sense?

&nbsp;&nbsp;&nbsp;&nbsp;Matthias

---

<div class="post-metadata">

**Author:** ![Siwek\_Jon](https://avatars.discourse-cdn.com/v4/letter/s/90db22/32.png) [@Siwek\_Jon](https://community.zeek.org/u/Siwek_Jon)\
**Post date:** [December 6, 2016, 8:43pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/11 "2016-12-06T20:43:06Z")

</div>

> "(v as T)$foobar” vs "as\<T\>(v)$foobar”.

Could just do some time trials to see which one people can type faster. There’s online typing speed tests that let you enter the test’s text. I consistently typed "(v as T)$foobar” faster.

> > Is there a reason "type(v)” can’t be stored?
> 
> Mind elaborating how you would expect to use that?

Was thinking I’d end up doing something like (more a personal habit/style thing):

&nbsp;&nbsp;local t = type(v);

&nbsp;&nbsp;if ( t == count )  
&nbsp;&nbsp;&nbsp;&nbsp;…  
&nbsp;&nbsp;else if ( t == int )  
&nbsp;&nbsp;&nbsp;&nbsp;...

But this example came to mind only because the draft didn’t have “v is T”. Point seems moot now.

> seems the existing switch could just infer that it's maching types by  
> looking at case values: if they are types, it'd use "is" for  
> comparision.

+1

> > - In the switch statement, I would require that a user either provide  
> > &nbsp;&nbsp;&nbsp;a default case or fully enumerates all cases. Otherwise it's too  
> > &nbsp;&nbsp;&nbsp;easy to cause harder-to-debug run-time errors.
> 
> Are we requiring "default" for the standard switch statement? If so, I  
> agree it makes sense to do the same here (otherwise, not so sure,  
> because of consistency).

The “default" case is optional. If it were required, I wouldn’t feel safer — maybe I don’t frequently make/see mistakes related to that and/or don’t find them hard to debug.

More often I’ve seen forgotten “break” at the end of a cases causing unintentional fallthroughs, but Bro does require explicit “break” or “fallthrough” statement to end cases.

- Jon

---

<div class="post-metadata">

**Author:** ![robin](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/robin/32/599_2.png) [@robin](https://community.zeek.org/u/robin)\
**Post date:** [December 8, 2016, 4:49pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/12 "2016-12-08T16:49:19Z")

</div>

Sounds like agreement on most suggestions, I'll update the web page  
shortly with the conclusions.

Couple further comments:

> I consistently typed "(v as T)$foobar” faster.

Personally, I see this more as a question of readability (as opposed  
to typeability :). But it's a matter of taste, and I'd be fine with  
using "as" instead of "cast\<\>".

> How is a timeout different from a failed computation due to a different  
> reason (e.g., connection failure, store backend error)?

I'm thinking such errors need to be addressed by the function itself,  
using whatever mechanism it deems appropriate (e.g., signaling trouble  
through its return value). The problem is that we don't have any  
further error handling mechanisms (e.g., exceptions) in the language  
right now (that's a project for some other time). I see timeouts as  
different as that's something the function itself may have a hard time  
catching internally, if at all; we couldn't rely on that. It's kind of  
last line of defense against trouble: if for whatever reason the  
function ends up never returning, we'll catch it and clean up at  
least.

> asynchronous operation. More generally, we need well-defined  
> semantics for composing asynchronous computations.

I think this would make sense only if we already had some model for  
propagating errors as part of the language. In the absence of that, I  
don't really see much to do here. For the normal (non-error) case,  
return values are just passed along as expected. And in the abnormal  
error case, there's not much else to do than abort the whole event  
handler.

Robin

---

<div class="post-metadata">

**Author:** ![Matthias\_Vallentin1](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/matthias_vallentin1/32/596_2.png) [@Matthias\_Vallentin1](https://community.zeek.org/u/Matthias_Vallentin1)\
**Post date:** [December 11, 2016, 11:49am UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/13 "2016-12-11T11:49:06Z")

</div>

> Personally, I see this more as a question of readability (as opposed  
> to typeability :). But it's a matter of taste, and I'd be fine with  
> using "as" instead of "cast\<\>".

Probably aligned with that thought is consistency and intuition: we  
don't have C++-style angle brackets in Bro, so "as" feels more in line  
with, e.g., type declarations of the form table[T] of U. Angle brackets  
would feel like a one-off, whereas space-separated short keywords fit in  
more naturally.

> I'm thinking such errors need to be addressed by the function itself,  
> using whatever mechanism it deems appropriate (e.g., signaling trouble  
> through its return value). The problem is that we don't have any  
> further error handling mechanisms (e.g., exceptions) in the language  
> right now (that's a project for some other time). I see timeouts as  
> different as that's something the function itself may have a hard time  
> catching internally, if at all; we couldn't rely on that. It's kind of  
> last line of defense against trouble: if for whatever reason the  
> function ends up never returning, we'll catch it and clean up at  
> least.

Indeed, timeouts are a last line of defense and generated by the  
runtime, as opposed to the function itself. You point out that these are  
two separate kinds of errors.

However, from a user perspective, I think errors should be treated  
uniformly. For example, when performing a Broker store lookup of a key,  
but the returned value has a type different than expected, script  
execution (in particular: a chain of asynchronous functions) should not  
continue. The same thing happens for a timeout, even though the  
runtime---as opposed to the actual function---generated the error.

The alternative strikes me as complicated, because it is application  
dependent. If a response to an asynchronous request doesn't arrive  
within a time bound, \*the effect\* might be equivalent to an error within  
the function itself. An invalid cast might generated a runtime error to  
the console (as Bro does now for use of uninitialized variables) and  
abort, whereas a timeout would simply abort.

In Bro, one way to represent represent errors would be as a record with  
an optional error field. If the field is set, an error took place and  
one can examine it. Either the function returns one of these explicitly,  
or the runtime (i.e., the interpreter) populates one.

Here's a strawman of what I mean:

&nbsp;&nbsp;&nbsp;&nbsp;type result = record {  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;error: failure &optional;  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;value: any &optional;  
&nbsp;&nbsp;&nbsp;&nbsp;}

&nbsp;&nbsp;&nbsp;&nbsp;function f(x: T) : result {  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;if (store lookup with bad cast)  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return [$error=..];  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;if (nothing to do)  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return ;  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;if (computed something)  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return [$data=42];  
&nbsp;&nbsp;&nbsp;&nbsp;}

&nbsp;&nbsp;&nbsp;&nbsp;local r = f(x);  
&nbsp;&nbsp;&nbsp;&nbsp;if (r$?error) {  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;// handle error, check if timeout or some other form of error  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;// occurred by inspecting the failure instance.  
&nbsp;&nbsp;&nbsp;&nbsp;}

Then, the runtime would allow for automatic chaining of asynchronous  
functions returning instances of type result:

&nbsp;&nbsp;&nbsp;&nbsp;function f(x: T) : result;  
&nbsp;&nbsp;&nbsp;&nbsp;function g(y: U) : result;

&nbsp;&nbsp;&nbsp;&nbsp;// Calls g iff f did not fail and return value unpacking succeeded.  
&nbsp;&nbsp;&nbsp;&nbsp;local r = g(f(x));

Here, the runtime could do the unpacking of f's result, checking for  
errors, and feeding the result data into g---assuming types work out.

In essence, this is a monad for a fixed set of types. The runtime  
performs the unboxing of return values automatically.

&nbsp;&nbsp;&nbsp;&nbsp;Matthias

---

<div class="post-metadata">

**Author:** ![Siwek\_Jon](https://avatars.discourse-cdn.com/v4/letter/s/90db22/32.png) [@Siwek\_Jon](https://community.zeek.org/u/Siwek_Jon)\
**Post date:** [December 11, 2016, 2:15pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/14 "2016-12-11T14:15:38Z")

</div>

> In Bro, one way to represent represent errors would be as a record with  
> an optional error field. If the field is set, an error took place and  
> one can examine it. Either the function returns one of these explicitly,  
> or the runtime (i.e., the interpreter) populates one.
> 
> Here's a strawman of what I mean:
> 
> &nbsp;&nbsp;&nbsp;type result = record {  
> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;error: failure &optional;  
> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;value: any &optional;  
> &nbsp;&nbsp;&nbsp;}

That type of structure for error handling/propagation seems fine to me.  
(Don’t see the advantage of using exceptions instead).

> Then, the runtime would allow for automatic chaining of asynchronous  
> functions returning instances of type result:
> 
> &nbsp;&nbsp;&nbsp;function f(x: T) : result;  
> &nbsp;&nbsp;&nbsp;function g(y: U) : result;
> 
> &nbsp;&nbsp;&nbsp;// Calls g iff f did not fail and return value unpacking succeeded.  
> &nbsp;&nbsp;&nbsp;local r = g(f(x));

I might prefer just doing the unpacking myself. Having separate/explicit checks for each individual return value would naturally occur over time anyway (e.g. when debugging, adding extra logging, improving error handling).

How common do you expect async function chains to occur? Any specific examples in mind where auto-chaining is very helpful? Maybe it’s more of a nice-to-have feature than a requirement?

- Jon

---

<div class="post-metadata">

**Author:** ![Matthias\_Vallentin1](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/matthias_vallentin1/32/596_2.png) [@Matthias\_Vallentin1](https://community.zeek.org/u/Matthias_Vallentin1)\
**Post date:** [December 13, 2016, 12:02pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/15 "2016-12-13T12:02:23Z")

</div>

> I might prefer just doing the unpacking myself. Having  
> separate/explicit checks for each individual return value would  
> naturally occur over time anyway (e.g. when debugging, adding extra  
> logging, improving error handling).

> How common do you expect async function chains to occur? Any specific  
> examples in mind where auto-chaining is very helpful? Maybe it’s more  
> of a nice-to-have feature than a requirement?

This is a bit of a chicken-and-egg problem: we don't have much use of  
when at the moment, because it's difficult to nest. If we provide a  
convenient means for composition from the get-go, I'd imagine we see a  
quicker adoption of the new features.

The standard workflow I anticipate is a lookup-compute-update cycle.  
Here's a half-baked example:

&nbsp;&nbsp;function test(attempts: count): result {  
&nbsp;&nbsp;&nbsp;&nbsp;if (attempts \> threshold)  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;NOTICE(...);  
&nbsp;&nbsp;&nbsp;&nbsp;return [$value = attempts + 1]; // no error  
&nbsp;&nbsp;}

&nbsp;&nbsp;function check(..) : result {  
&nbsp;&nbsp;&nbsp;&nbsp;local key = fmt("/outbound/%s", host);  
&nbsp;&nbsp;&nbsp;&nbsp;local r = async lookup(store, key);  
&nbsp;&nbsp;&nbsp;&nbsp;if (r?$error)  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return local?$error;  
&nbsp;&nbsp;&nbsp;&nbsp;return async put(store, key, failed\_attempts + 1);  
&nbsp;&nbsp;}

With the proposed monad, you could factor the implementation of the  
check function check into a single line:

&nbsp;&nbsp;local r = put(store, key, test(lookup(store, key)));  
&nbsp;&nbsp;  
Conceptually, this is equivalent to "binding" your function calls:

&nbsp;&nbsp;local r = lookup(store, key))  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;\>\>= test  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;\>\>= function(x: count) { return put(store, key, x); }; // curry

...where \>\>= would be the bind operator performing error/value  
unpacking. Less functional clutter with the "do" notation:

&nbsp;&nbsp;do  
&nbsp;&nbsp;&nbsp;&nbsp;x1 \<- lookup store key  
&nbsp;&nbsp;&nbsp;&nbsp;x2 \<- test x1  
&nbsp;&nbsp;&nbsp;&nbsp;x3 \<- put store key x2

The last two snippets derail the conversation a bit---just to show where  
these ideas come from.

Note that the function "test" is synchronous in my example and that this  
doesn't matter. The proposed error handling makes it straight-forward to  
mix the two invocation styles.

&nbsp;&nbsp;&nbsp;&nbsp;Matthias

---

<div class="post-metadata">

**Author:** ![Siwek\_Jon](https://avatars.discourse-cdn.com/v4/letter/s/90db22/32.png) [@Siwek\_Jon](https://community.zeek.org/u/Siwek_Jon)\
**Post date:** [December 13, 2016, 4:53pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/16 "2016-12-13T16:53:41Z")

</div>

For prototyping purposes, I see the convenience in that, but wonder if the runtime will do something that’s useful and widely applicable enough for that to translate directly into production code. What exactly does the runtime do if “lookup” fails here besides turn the outer function calls into no-ops?

Just guessing, but in many cases one would additionally want to log a warning and others where they even want to schedule the operation to retry at a later time. i.e. the treatment of the failure varies with the context. Is this type of composition flexible enough for that?

- Jon

---

<div class="post-metadata">

**Author:** ![Matthias\_Vallentin1](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/matthias_vallentin1/32/596_2.png) [@Matthias\_Vallentin1](https://community.zeek.org/u/Matthias_Vallentin1)\
**Post date:** [December 13, 2016, 5:42pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/17 "2016-12-13T17:42:10Z")

</div>

> \> local r = put(store, key, test(lookup(store, key)));
> 
> For prototyping purposes, I see the convenience in that, but wonder if  
> the runtime will do something that’s useful and widely applicable  
> enough for that to translate directly into production code. What  
> exactly does the runtime do if “lookup” fails here besides turn the  
> outer function calls into no-ops?

The runtime doesn't do anything but error propagation. That's a plus in  
my opinion, because it's simple to understand and efficient to  
implement.

> Just guessing, but in many cases one would additionally want to log a  
> warning and others where they even want to schedule the operation to  
> retry at a later time. i.e. the treatment of the failure varies with  
> the context. Is this type of composition flexible enough for that?

It's up to the user to check the result variable (here: r) and decide  
what to do: abort, retry, continue, or report an error. Based on what  
constitutes a self-contained unit in an algorithm, there are natural  
points where one would try again. In the above example, perhaps at the  
granularity of that put-test-lookup chain. To try again after a timeout,  
I would imagine it as follows:

&nbsp;&nbsp;&nbsp;&nbsp;local algorithm = function() : result {  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return put(store, key, test(lookup(store, key))) &timeout = 3 sec;  
&nbsp;&nbsp;&nbsp;&nbsp;};

&nbsp;&nbsp;&nbsp;&nbsp;local r = result();  
&nbsp;&nbsp;&nbsp;&nbsp;do {  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;r = algorithm();  
&nbsp;&nbsp;&nbsp;&nbsp;} while (r?$error && r$error == timeout);

Based on how flexible we design the type in r$error, it can represent  
all sorts of errors. The runtime could abort for critical errors, but  
give control back to the user when it's a recoverable one.

&nbsp;&nbsp;&nbsp;&nbsp;Matthias

---

<div class="post-metadata">

**Author:** ![robin](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/robin/32/599_2.png) [@robin](https://community.zeek.org/u/robin)\
**Post date:** [December 13, 2016, 6:52pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/18 "2016-12-13T18:52:55Z")

</div>

> \> type result = record {  
> \> error: failure &optional;  
> \> value: any &optional;  
> \> }

I don't really like using a record like that, as that would associate  
specific semantics with what's really a user-definable type. I.e.,  
we'd hardcode into the language that the record used here needs to  
have exactly these fields. On top of that, it also feels rather clumsy  
to me as well. Indeed it's one of the things I don't like about the  
current Broker framework API (which currently has to do it that way  
just because there's no better mechanism). My draft proposal extends  
opaque types to support conversion to boolean to allow for more  
elegant error checks. We could generalize that to support other data  
types as well, although I don't really see a need for going beyond  
opaque right now (assuming we also add the cast operator, per the  
proposal, so that we aren't passing anys around).

> With the proposed monad, you could factor the implementation of the  
> check function check into a single line:
> 
> &nbsp;&nbsp;local r = put(store, key, test(lookup(store, key)));

Honestly, I think this operates on a level that's quite beyond any  
other concepts that the language currently offers, and therefore  
doesn't really fit in. I'm pretty certain this wouldn't really enter  
the "active working set" of pretty much any Bro user out there. For  
the time being at least, explicit error checking seems good/convinient  
enough to me.

> r$error == timeout

> Based on how flexible we design the type in r$error, it can represent  
> all sorts of errors.

That's a piece that I like: Being able to flag different error types.  
Maybe the conversion to bool that I proposed originally should really  
be a conversion to a dedicated error type, so that one can  
differentiate what happened. I need to think about it, but that could  
eliminate the need for simply aborting the event handler on timeout  
(which in turn would help with the problem that Bro isn't good at  
aborting code)

Robin

---

<div class="post-metadata">

**Author:** ![Matthias\_Vallentin1](https://yyz1.discourse-cdn.com/flex011/user_avatar/community.zeek.org/matthias_vallentin1/32/596_2.png) [@Matthias\_Vallentin1](https://community.zeek.org/u/Matthias_Vallentin1)\
**Post date:** [December 13, 2016, 7:51pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/19 "2016-12-13T19:51:39Z")

</div>

> I don't really like using a record like that, as that would associate  
> specific semantics with what's really a user-definable type.

It was only meant to illustrate the idea of error handling and function  
composition. These ideas still hold up when substituting the  
user-defined result type with the proper language-level construct.

> We could generalize that to support other data types as well, although  
> I don't really see a need for going beyond opaque right now

Going beyond opaque would have the advantage of applying a uniform error  
handling strategy to both synchronous and asynchronous code. That's not  
needed right now, because "when" doesn't make it easy to handle errors,  
but it could be an opportunity to provide a unified mechanism for a  
currently neglected aspect of the Bro language.

> Maybe the conversion to bool that I proposed originally should really  
> be a conversion to a dedicated error type, so that one can  
> differentiate what happened.

I like that.

&nbsp;&nbsp;&nbsp;&nbsp;Matthias

---

<div class="post-metadata">

**Author:** ![Seth\_Hall3](https://avatars.discourse-cdn.com/v4/letter/s/d6d6ee/32.png) [@Seth\_Hall3](https://community.zeek.org/u/Seth_Hall3)\
**Post date:** [December 14, 2016, 2:39pm UTC](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526/20 "2016-12-14T14:39:16Z")

</div>

I like that too. Having nicely generalized error handling in Bro would be such a huge benefit for script authors.

.Seth

[Next page](https://community.zeek.org/t/proposal-language-extensions-for-better-broker-support/4526.md?page=2)
