Showing posts with label software. Show all posts
Showing posts with label software. Show all posts

Monday, 16 March 2015

Overthinking Scala Parsers

Let me start by telling how much I love Scala's Parsing Combinators framework. I had a nice experience some time ago, working with tools based on very similar ideas on Haskell, and I felt glad to realize that the Scala guys where able to fully maintain the power and simplicity of functional parser combinators while, at the same time, integrate it so gracefully to the object oriented world and Scala's complex type system.

Anyway, I have no intention to write about the framework itself (It's very well documented and it's not hard to find lots of tutorials and examples), but rather focus on sharing our recent experience dealing with some of it's rough edges. There where mainly two particular subjects that, though subtle, I found interesting. One, once you got your parser to work, what can you do to make it nicer? And two, how do you test it?

Imagine you have some code like this one:

object AwesomeParser extends RegexParsers {
  val foo: Parser[Foo] = ... // Some foo parser
  val bar: Parser[Bar] = ... // Some bar parser
  val yetMoreParsers = ...
}

For the sake of the post, let's assume that this is a complete, working example. We could also give for granted that the code is very clean and we attend all the common functional good practices such as avoiding side effects and defining larger parsers as the combination of very small and simple ones. Ok, so the logic is working, and the code is clear; what could possibly be wrong?

Well... Nothing is "wrong" I guess... But there are a couple of things here that I felt kind of unpleasant. For example, let's say that the day comes that I have to add a new parsers defined as one or more foo and one or more bar.

object AwesomeParser extends RegexParsers {
  val manyFoos = foo.+
  val foo: Parser[Foo] = ... // Some foo parser
  val bar: Parser[Bar] = ... // Some bar parser
  val manyBars = bar.+
  val yetMoreParsers = ...
}

This may still look ok, it will even compile just fine, but truth is the manyFoos implementation is broken. Why? Because it references a value not yet initialized. At this point you may be thinking "Oh, God! Just put the manyFoos val at the bottom, you obsessive freak!" and, sure, that would fix this scenario, but once you have a good fifty something parsers with combinations and recursions finding the right spot to place your parsers can be tricky. Besides I don't want to worry about declaring stuff before using it (What am I, a caveman?).
We could replace all those vals with defs, but that would also introduce some unnecessary overhead, since parsers would have to be combined every single time they are referenced... I find that the most satisfying alternative (in my experience, of course) is to define parsers as lazy initialized values. That way the parsers would get combined one single time, but you can pretty much place them wherever you want to.

object AwesomeParser extends RegexParsers {
  lazy val manyFoos = foo.+
  lazy val foo: Parser[Foo] = ... // Some foo parser
  lazy val bar: Parser[Bar] = ... // Some bar parser
  lazy val manyBars = bar.+
  lazy val yetMoreParsers = ...
}

Another annoying thing about this implementation is how dirty the interface is. Sure, the code may look clean and easy now, but wait until you get that AwesomeParser object and press ctrl+space...
As you may notice, my AwesomeParser object extends the RegexParsers trait (though it may have been any other member of the  Parsers subtrait family). These traits pretty much define little parsers ecosystems, along with the types and functions needed to manipulate them. Problem is, the amount of methods and types defined is HUGE (and all the parsers we added don't really help the problem). Also, there is no standar entry point. A RegexParser may be executed using the parse or parseAll methods, while a Json parser provides parseFull and parseRaw.
Adding all up, it is quite bothersome for the parser user to navigate this large interfaces looking for the way to execute the parser and ignoring amost everything else. And even when they find the right parse method, most of it's versions require the user to pass along the parser that should be used to parse. For example, our AwesomeParser may be called like this:

AwesomeParser.parse(AwesomeParser.manyFoos,"foofoofoofoo")

Now, lets face it. It is great to have the possibility to chose the parse approach, but most of the time you don't need that many options. On those cases, a simple way to get a cleaner interface for your parser may be to "hide" it's definition in an inner declaration and just expose the messages you want.

object AwesomeParser {

  def apply(input: String) = Definition.parse(Definition.manyFoos, input)

  protected object Definition extends RegexParsers {
    lazy val manyFoos = foo.+
    lazy val foo: Parser[Foo] = ... // Some foo parser
    lazy val bar: Parser[Bar] = ... // Some bar parser
    lazy val manyBars = bar.+
    lazy val yetMoreParsers = ...
  }
}

One last thing to object about this implementation is that it is pretty much impossible to extend. A singleton object may seem like a great implementation for a parser; it grants global access and probably leads to a healthy stateless implementation, but what if somebody wants to extend it or make a subtle variation on the grammar? Well, why can't we have booth? It is as simple as offering a trait implementation of your parser and make a default singleton instance to extend it. That way you keep all the perks of a global, well-known parser without sacrificing extensibility.

object AwesomeParser {

  def apply(input: String) = Definition.parse(Definition.manyFoos, input)

  protected object Definition extends AwesomeParserDefinition
  trait AwesomeParserDefinition extends RegexParsers {
    lazy val manyFoos = foo.+
    lazy val foo: Parser[Foo] = ... // Some foo parser
    lazy val bar: Parser[Bar] = ... // Some bar parser
    lazy val manyBars = bar.+
    lazy val yetMoreParsers = ...
  }
}

Now lets talk testing. I can't even imagine developing/maintaining a parser without a full set of tests. The thing is that, in my experience, parser tests tend to be much MUCH larger than the parsers themselves and require lots of attention and care. The simpler and more expresive they are, the easier it will be to implement changes to the parsed grammar. That's why we packed a couple of simple tools we developed to improve the parser testing and published them here for anyone to use.

And that's all I have to say about that.


N.

Wednesday, 5 November 2014

Travis, o cómo elegir un CI para tu proyecto Java open source

Recientemente me incorporé al equipo de desarrollo de Arena, nuestro framework de MVVM UI pensado para que alumnos de diversas carreras universitarias den sus primeros pasos en el desarrollo de interfases de usuario. Hicimos un sprint para empezar a familiarizarnos con las herramientas y ahí detectamos algunos problemas en la forma de trabajo, que dificultaban un poco la incorporación de nuevos desarrolladores.
Los inconvenientes principales eran:
  • Para trabajar en un proyecto (hoy Arena se compone de 7 proyectos distintos) no había otra opción que compilar todos los demás, ya que al no estar sistematizado el deploy no siempre el SNAPSHOT se correspondía con la última versión del código. Entonces perdía un poco de sentido la separación en artefactos distintos.
  • Algunas veces se pusheaba al repositorio código que rompía los tests, o incluso no compilaba. Al no tener releases estables, podía darse el caso de que un deploy del SNAPSHOT rompiera los trabajos prácticos de los alumnos de una cursada, como sucedió alguna que otra vez.
  • Los artefactos se desplegaban en un repositorio propio y no en Maven Central, lo que implicaba un paso extra de configuración a la hora de preparar el entorno de desarrollo. Trataremos este tema en el siguiente post de esta serie.
¿Cómo solucionar esto? Automatizando estas tareas repetitivas por medio de un servidor de integración continua. No es el propósito de este artículo discutir los beneficios de esta práctica (para eso puede leerse este artículo de Martin Fowler), así que vamos a pasar directamente a la solución que puse en marcha.

Eligiendo un CI para Arena

Para los ansiosos, la implementación final de lo que contaré a continuación (incluyendo el próximo post) puede verse en el repositorio de Arena en GitHub. De todos modos, recomiendo leer cada una de las decisiones que tomé para entender qué es lo que hace el CI cada vez que se sube código nuevo al repositorio.

CIs Candidatos

Además de que fuera gratuito para Open Source Software (OSS), se necesitaba que cumpliera tres características:
  1. que fuera fácil de configurar y que no necesitara hosting propio.
  2. que soportara Java de forma nativa.
  3. que tuviera algún mecanismo para subir el settings.xml, necesario para poder hacer el deploy a Sonatype.
Entonces, en ese orden, descarté estos 3 productos:
  1. Jenkins, quizás el más conocido y flexible de los que voy a mencionar, pero con el overhead de tener que hostearlo en servidor propio y configurar todos los plugins.
  2. Semaphore, que suele ser mi opción preferida, pero como todavía no soporta Java “out of the box” es necesario bajar Maven en cada build.
  3. Y por último, drone.io, donde no encontré cómo solucionar el problema del settings.xml.
Finalmente mi elección fue Travis, obviamente en su versión para OSS.
(Un comentario al margen: Travis tiene también una versión para repositorios privados y, gracias a un acuerdo con GitHub, la ofrecen de manera gratuita e ilimitada para estudiantes mediante el student developer pack.)

Configurando Travis

La configuración de Travis está basada en el principio de convention over configuration. Para empezar a compilar y correr nuestros tests, basta crear un archivo llamado .travis.yml e indicar en su interior en qué lenguaje está realizado nuestro proyecto, en nuestro caso Java. El primer archivo de configuración se veía así:
language: java
¿En serio? ¿ya está? No del todo, pero ya estamos en condiciones de probar que funciona avisándole a Travis que monitoree el proyecto (esto se hace desde tu perfil) y luego haciendo un git push de algún cambio (aunque una mejor idea sería crear una branch de prueba y hacer un pull request).
Por defecto, Travis empezará a monitorear el repositorio y disparará un build por cada commit que se suba a cualquiera de sus branches, incluyendo pull requests. Una vez concluido dicho build, se encargará de notificar a GitHub el resultado, quien a su vez nos lo mostrará a nosotros como se ve en la siguiente imagen:
Pull requests validados por Travis
Con esa simple línea de configuración que le dimos, lo que se ejecuta es el default para Java, que consta de dos pasos:
  1. mvn install -DskipTests=true, baja todas las dependencias y compila el proyecto. ¿Y para qué le pone ese flag si también queremos que corra los tests? Cuando algo anda mal, Travis diferencia entre un build errored y uno failed, donde el primer estado está asociado con un error en la instalación (no compila, no se encontró alguna dependencia) y el segundo con alguna validación que se ejecuta (normalmente un test fallido). Entonces si llegara a fallar la instalación, ni siquiera se corren los tests y sabemos con certeza que el resultado es errored
  2. mvn test (no necesita explicación)
Hasta acá llega el primer post de esta duología. En el siguiente post explicaré los pasos necesarios para publicar la aplicación en el Maven Central Repository, utilizando Sonatype OSS como intermediario.
¡No dejen de compartir sus proyectos, ideas o experiencias en los comentarios!