sábado, 18 de abril de 2015

How Users Adopt Products

In one of the companies that I used to work for, a group of very smart engineers decided that we needed to opt for a unified development tool for all developers. There were two tools in the market: product corpX and product rebelY. Since they had invested some time and plugins to refine corpX, they wanted to force its use. They had little success on adoption rates. They claimed that some of the problem was in those badass users of rebelY that kept convincing the rest of the team to shift. These badass users were evangelizing the features, they were forming groups of users, they were showing off their results, and even they were wearing proudly political propaganda of the rebelY product (wearing t-shirts and other pride items). Those rebels were resisting and were winning the battle. So they decided it was a moment to force the good in them.

The reaction in itself seemed to me motivated more by emotion than reason (or data), so I decided to write what little I know about product adoption because my friends were going in the wrong direction.


A key feature to highlight is desirability; or said in other terms, that thing that we want and don’t have. Desirability is what drives people to product love. And product love is blind and passionate. It makes users make comments about how they feel empowered by the product. And eventually, those comments reach us, our friends, family members or colleagues. “Can’t believe Marco can do that in one click”. Marco is cool we think. I need to get it (the tool not Marco).


And yes, we really believe more what those close friends say or do than any other corporate advertisement or political propaganda. So what inspires people to make those comments? The answer is simple: Users do not evangelize to their friends because they like the product, but because they like their friends (and the company they work for) -- to quote Kathy Sierra. Users enjoy the bigger context not the product. The bigger context is to create good software and be productive. And results tell the story: "I can create a test in a fly!" "See how Marco can search the repo!" "I can catch bugs before they happen". All of that makes me a better developer so I want to share that with my friends.


The problem is that these badass users are so empowered that they do not shut up. They continue talking. But what is worse is that they do not need to talk. People see the way they work and the results that they are getting. That is even more powerful. They are getting better results.

My claim is not to treat people like puppies. We are rational people making conscious choices. We like to have tools that empower us. And developers need to always think in the big context. The result of the initiative, as in any European movie, is left for you to figure out.

domingo, 13 de abril de 2014

Convert a JAR into Bundle

Personal notes on how to quickly create a bundle based on a list of jars or other dependencies. It uses the content-package and embedded maven configuration to easily deploy the bundle.

Normal and Painful Way
The most difficult way is to follow the standard way and create a manifest file with all the description available. You will also need to figure out the dependencies of your bundle. This blog will show the easy way to package a bundle with all the dependencies that you need.

Step#1: Easy Way

The project is quickly created by the classic mvn archetype. The idea is to follow the structure. 

myprojectlib/depency-pkg-libs

where myprojectlib is the root project and depency-pkg-libs the module with the declared dependencies. We will not include any other code into the folders. Also for the purposes of example, we will use the JUnit Sling Testing on the server side. This bundle is not included in the default distribution and requires JUnit as a dependency.

Step #2: Create the First module

mvn archetype:generate

Make sure you create a basic maven repository with the pom project as root.

In our case the  myprojectlib will look like:

<packaging>pom</packaging>

Step#3: Create the  Content-package
Create a package depency-pkg-libs inside the root directory. Make sure you change the packaging so that it becomes content-package

<packaging>content-package</packaging>

Step#4: Add Embedded List
To your list of build plugins, add the content-package-maven plugin

<groupId>com.day.jcr.vault</groupId>
<artifactId>content-package-maven-plugin</artifactId>

In the list of embedded element add those dependencies that you need. In my case, the junit.remote:

<embeddeds>
  <embedded>
     <groupId>org.apache.sling</groupId>
      <artifactId>org.apache.sling.junit.remote</artifactId>
     <target>${package.root}/install</target>
  </embedded>
</embeddeds>

You should also add the same artifacts listed in embedded to your list of dependencies. Otherwise those will not be installed.

     <dependency>
          <groupId>org.apache.sling</groupId>
          <artifactId>org.apache.sling.junit.remote</artifactId>
          <version>1.0.6</version>
          <scope>provided</scope>
      </dependency>

Actual Code
You can check the actual code from bitbucket. The meat is in the pom.
References
This blog is based on the following resources.


lunes, 7 de abril de 2014

Retrieving Service from Request

Quick way to retrieve a service from a slingRequest

private <E> E getServiceFormRequest(
final SlingHttpServletRequest request, Class<E> serviceType ){
        final SlingBindings bindings = 

(SlingBindings) request.getAttribute(SlingBindings.class.getName());
        SlingScriptHelper slingScriptHelper = bindings.getSling();
        return slingScriptHelper.getService(serviceType);
    }

domingo, 23 de marzo de 2014

Date Range Query Medicine for CQ AEM

One of the most annoying features of xpath queries is the performance problem of date range queries. QueryBuilder tends to perform badly. You can see Marcel's presentation in AdaptTo if you want to know the details. Thus, when faced with this problem, there are several alternatives:
  • modifying the granularity of the dates,so that you select up to the day and exclude hours/min/secs
  • using a node structure in which the date is saved as the path; for example,
    • yyyy/mm/dd/my_content_node
  • or converting the date property into a long 
I am going to focus on the third one: ways to convert a date property into a long property.

Approach
The basic idea is to leverage the JCR Observation offered in AEM to add a property of type long each time a property of type date is created or modified.

First I created my own namespace (thanks this notes) by just going to

http://<host>:<port>/crx/explorer/nodetypes/index.jsp

I selected my own namespace netval so that there will be no conflicts with existing properties. 

Then I registered my class DatePropertyEventListener that implements the EventListener interface. The class needs to use the @Component annotation. You have to register for the event listener in the @Activate method. Make sure you get the ComponentContext in the activate method in order to register for the event. 

In my onEvent I used the EventIterator to get the property. I then filtered those properties that are of type date. Conveniently the Property class can return almost any type of value being String, long, Calendar, etc. In my case I needed the value as a long. 

Once I have the value, I can set it in the new property that has netval as the namespace and the name of the property in order to avoid collisions. For example, if the attribute name was activationDate the new attribute would be:

netval:activationDate

Code
The first thing is to register and de-register the event listener in the @Activate and @Deactivate methods. 

A quick look at my DatePropertyEventListener class reveals the following. The observationManager and admin session are class attributes. They will be used in the @Deactivate  nd Activate methods. It is in the @Activate method where we register the event listener. 

observationManager = adminSession.getWorkspace().getObservationManager();
observationManager.addEventListener(this,
Event.PROPERTY_ADDED|Event.PROPERTY_CHANGED,
                "/content/news/", // absolute path
                true, // isDeep
                null, // uuid
                null, // nodeTypeName
                true // no local
        );

Note that the absolute path is restrictive to the root path of your content. I highly recommend being as specific as possible in order not to receive too much noise.

The onEvent it is easy to get the date property

        while( eventIterator.hasNext() ){
            final Event event = eventIterator.nextEvent();
            if(PropertyType.DATE == getPropertyType(event) ){

Caveats and Warnings
This procedure should be disabled for operations that modify many nodes, like during a migration of content.

A suggestion is to have a producer-consumer data structure (like LinkedBlockedQueue) that will allow you to process the events more efficiently.

lunes, 18 de marzo de 2013

Heavy duty ETL: Extraction Transformation and Load

This is a cross post blog from Cloud Data Viz project

These days, many people are talking about Big Data. However, very few talk about how big is Big Data and about all the different components that need to be considered prior, during, and after running a system based on Big Data. Many don't know the time it takes to extract the data. Others forget that they need to validate it. And only a few mention the loading tools they use to speed up the process of loading millions of rows into a system (not to mention normalizing the data while you are loading it). Therefore, I thought our "little" project could be of interest. This post will explain all the different issues you need to consider before implementing a system with Big Data.

Data Volumes
The amount of data that we use for the Climate Viz project is astronomical. Just to give an idea, we are collecting temperature data from satellites at a 0.25 resolution, which means that there are 864,000 data points being collected every 3 hours. At the end of just one day, we get close to 7 million data points. Just for one day! And we haven't mentioned transformation and other data aggregation statistics that we need for the project.

Extraction
Every project always comes with dirty work. And extracting and loading the files is the dirty part of this project. Visualizing the data in maps and writing the UI controls is the easy part. But getting all the data in and validating it, that's where the pain begins.

We use a two-step process for extracting the files from NASA Giovanni. Setting up a wget script that downloads the GRIB files was easy. Then we process them with pygrib and slice them in order to have them ready for GAE.

For data is never going to change, we generate master tables. In our case, latitude and longitude are fixed for every resolution and can be generated programatically. Two for loops ( x and  y ) and you think you are done but, when you are dealing with Big Data, that translates into several troubles. First consider that the loops could take so long that you could reach timeouts as part of the request and post methods of the web layer. Then, the same thing could happen even if you make a dummy get method to generate all the different tasks using the queue.

Validation
Once you start dealing with data collected from different sources, the first need is to validate them. Most likely you would like to have a visualization tool that compares and contrasts your data. Unfortunately, you can't use excel because it has a limit of 65,535 rows (far too limited for big data sources).

Thus, we are left with building our own tools to validate our data. Think about it because it makes a big difference. Google Maps was a useful interface for us. Also, Google drive is another option for loading big files.

Transfer 
Of course, transferring data from your source to the destination is another issue. If you have a normal ETL tool, then your problem is solved. However, the rest of us have to deal with several issues.  For each POST, there is a limit in both the size of the file we are sending as well as the time it takes to process it. I had to play around with how fast GAE was processing the files. I started with 15000 lines per file and had to go down to 200 lines. Otherwise, I received a timeout error in the operation for 5K, 2K, and 800. Luckily, I was generating files with a test script but at times is not easy to generate data.

I also generated a script to scan the files and post them automatically (using multi-part). Even with a queue of 40 tasks a second, the time to process the files is agonisingly slow (5 seconds per 100 lines) and it could probably take several months to load. :(

Another trick that I used was to slice the files to upload.  Each file contained a line order that will help me aggregate the data later.

Conclusion
I hope the ideas exposed in this post will help you in your Big Data projects.

miércoles, 31 de octubre de 2012

Delegando Tareas con App Engine Task Queues

Contenido de la presentación para el Google Dev Fest en Barcelona (9/11/2012)

Ver otras ponencias

Título:
Dejando tareas a Google App Engine: Introducción a Task Queues

Presentacion: https://www.box.com/files/0/item/f_3909761486

Contenido de la Presentación:

Introducción
  • Concepto General: herramienta para los procesos "background"
    • Creado por el limite de 30 segundos por peticion a 10 minutos
  • Visión general del API: los dos tipos de queues: push vs pull
  •  
  • Casos de Uso
    • Migración de Datos
    • Procesos Batch
    • Sincronización Avanzada
    • Integración con crons
Uso / Desarrollo / Cómo Dividir Tareas
  • Conceptos generales para dividir y conquistar (Divide and conquer)
  • Framework de Google y su Parametrización de las tareas
  • Cómo llamar a los "workers" y consideraciones especiales
    • idempotency: qué es
    • consideración para el entorno "cloud"
Configuración y Control
  • Técnicas de afinamiento
  • Repasando los atributos de una queue
    • frecuency/rate
    • bucket size
    • etc
  • Tu amigo el "control panel"
  • Versionado
Mejores Prácticas
  • Cómo diseñar para parar en "masa"
  • Control de tareas
Otros Frameworks y Herramientas
  • Map Reduce
  • appengine-pipeline

sábado, 18 de agosto de 2012

Analyzing Web Performance: love my RUM

It was news to me the other day while watching my favorite tv channel (Google Developers on YouTube) that I ran into this presentation about web performance. Previously I had no idea that modern browsers could measure metrics like "dns lookup", "server connection time", "server response time". With the awesome work from the W3C Performance Working Group, now this is a reality.

Added to the new and latests trends in continuous build cycle and automated testing, now there is the fact that you can get your RUM (Real User Metrics) as part of Google Analytics. There is always the need to do your performance testing but you can also see how the site evolves and how you could even monitor in real time what is going on. So I went I checked the Google Analytics and indeed, they are providing with all this data.

I remember a long while ago when a client had some problems with performance and we were working in the dark until we found out it was the dns lookup time! Imagine how much easy it would have been had we used this data.

For the curious I am posting the image of the lifecycle for a request. All this data could be obtained from the "Developers tool" from most modern browsers. Safari is still lagging.