Showing posts with label data science. Show all posts
Showing posts with label data science. Show all posts

Tuesday, July 9, 2019

The Book of Job

Rumors had been whispering for a month. The new financing round hadn’t gone well, the expected deals hadn’t materialized, and something had to give. Free lunches had been cancelled, the weekly company updates had changed to biweekly and the last one of these had been mysteriously cancelled. I started updating my resume.

One Tuesday I learned from backchannels that there would be “a layoff of not insignificant size” the next day. Despite some reassurances from a colleague that I should be safe, that night I tossed and turned until 3am, when I firmly concluded that I would be laid off. A mix of terror and exasperation hit me. I’d been laid off just a year previously, and I didn’t want to be unemployed again. When I awoke Wednesday morning, the commute to work felt like marching to the gallows. 

The process was efficient. They called half the company one-by-one into private offices, handed us exit papers and collected our badges and laptops. By 10am I was out, and by 11 had rendezvoused with other ex-employees at the Friendly Toast, where the bar was empty but open. The ensuing hours reached a level of day drinking rivaling my senior year St Patrick’s Day.

When I sobered up, I was sad, bitter, exhausted but excited. I was sad because CiBO had been my most enjoyable job. I had done interesting work with smart coworkers serving a great high level mission. It paid well, hadn’t been too stressful, and I could run to work on occasion. It had even taken me on a crazy 1 day trip to Malaysia. While the job didn’t trap me in the office long, I found myself home practicing Scala and studying the growth stages of corn. It bristles me now how fruitless this effort feels. Furthermore, the immediate termination was much rougher than the 1 month notice GE had given me. I had no time to mentally prepare to wake up the next day with absolutely nothing to do. I was bitter that after this long journey of changing careers, having spent so much time reflecting on what I wanted and then working so hard to actually get there, I had ended up with nothing. Twice. And during winter again. My browser cookies still remembered the Massachusetts unemployment website. I joked that I was now an expert in company collapses, of all different sizes. I got plenty of sympathy laughs, but when faced with the reality of yet another job hunt, I was exhausted before I even began.

Considering how much I care about my career, it’s a bit ironic that I’ve spent so much time funemployed that I'm able to name each period. Leaving Hong Kong, backpacking around Asia and returning to the US was my SabbatiCal. The period between GE and CiBO that included two international trips were my Callivanting days. This period? More like a Calamity. While I’m lucky to have had so many employment breaks - so many of my friends haven’t had any - this one was ill-timed and unwelcome.

However, I was excited because I had options for in 2019, data scientists are in short supply.  CiBO had been a fantastic tech environment where I’d worked closely with great software engineers. I had accrued enough confidence that I was a pretty badass data scientist and almost immediately began working with 20+ recruiters. I quickly realized though that I wanted, and had enough savings, to take my time. I wanted to explore transitioning out of a technical role, perhaps into strategy or product management. I wanted to return closer to the energy sustainability domain. And I wanted to move back to Asia. It was a tough multiple-criteria decision problem to optimize.

The best part of this Calamity was the many people who reached out to me and helped. It seems like I caught up with 100 friends that first week, juggling all time zones to the east and west. I chatted with friends about their professional lives and gained valuable insights into how other jobs worked. I had plenty of deep conversations that convinced me that my heart was still in Asia. I specifically targeted Beijing, Shenzhen, and Saigon, cities where I felt I could find cool jobs and cool people.

Saigon had vibed with me when I first visited during the SabbatiCal. I knew there was a decent tech scene, with a large concentration of foreign “digital nomads” utilizing a local ecosystem full of talented (and cheap) coders. I wasn’t sure what the options were for someone with no local country or language background, but my Saigon-dwelling friend Sam Axelrod connected me to someone who’d know. This guy gave me a rundown of the work international consultancies were doing, the locally-disconnected digital nomad scene, and the rapid government-backed digitization across the economy. He inadvertently went on a rant against those big name consultancies collaborating with government officials and multinational corporations to perpetuate modern colonialism. Having also lived in an Asian former colony, this rant won me over - he had expressed my views, albeit much more profoundly and eloquently. When he told me that in his previous role leading a UN bureau, he had made all his employees learn Vietnamese, I was sold. Then the conversation took a turn. “I lead a startup consultancy now. We have a Taiwanese manufacturing client and only one Chinese speaker on staff…and all of our work is really about using data to drive decisions….actually we could really use someone like you.” And so the informational chat turned into a job interview. A week later, I booked tickets to Ho Chi Minh City.
This is a fine ad for funemployment


In the meantime, I had already had a trip to New Orleans to visit my friend Jason Siu and partake in Mardi Gras. I took my employment frustrations out on hurricanes and Sazeracs, and somehow found myself walking down 10 blocks of Bourbon Street double fisting beers looking for Jason. The next morning, I awoke wearing a bushel of beads and needing to dry heave. I had scheduled a handful of recruiter calls before a late afternoon flight back to Boston. As I laid down on the couch in utter pain, I talked to Amazon on speaker phone and tried to go through my work history. I didn’t get a second interview. I barely made it to the airport, where I passed out on the dirty floor while JetBlue delayed us for 2 hours. When I took my middle seat, the old man sitting window asked me, in a volume indicating he was hard of hearing, “Did you enjoy the parades?” I did my best not to puke on him.

Back home, I planned an Asia trip to be part fact-finding mission and part friends catchup tour. I eschewed traveling to new places in favor of looking for jobs in familiar cities. I initially outlined a Saigon to Hong Kong to Shenzhen to Beijing to Paris to London trip, allotting myself 3 weeks. When my friend Doug Heimburger sold me hard on his 40th birthday celebration, I swapped out Shenzhen for Tokyo, then dropped Paris. I realized my dates in Hong Kong would coincide with Tosscars, the annual awards ceremony/party for the Hong Kong ultimate community. The ceremony’s hosts are secret until the event itself, and I had never been a host. I texted the organizers, and asked them if they wanted a super secret host. They replied that the theme was Carnival, and asked me to bring over Mardi Gras party supplies. Coincidentally I got that text while in a cafe in New Orleans, and walked outside to see a street vendor hawking party jackets. I got a sympathy $20 unemployment discount, and that jacket has proven to be one of my best investments.
I made a pun so bad, Vietnam decided to banh mi

Even since 2016, Saigon’s change was noticeable to me. There were more foreigners around District 1, Southeast Asia’s tallest skyscraper on the horizon and flat whites served in some coffeeshops in District 3. My first sight was a continuous stream of motorbikes street with no crosswalks and I had to relearn street crossing in Vietnam (with confidence and without eye contact). 

The startup consultancy was located above a clothing store and consisted of 6 employees. Though I’d be the only non-Vietnamese speaker, I knew I’d fit in well. My main worry with the role was whether I would stagnate technically. Though I was excited to learn about the Vietnamese economy and management consultancy in general, there was a good chance that the clients wouldn’t be ready for any interesting modeling, and I wasn’t sure I was ready to give that up. On my last day there, I was offered a role as an analyst, with the expectation that if I proved I could adapt to Vietnam, I would create and lead the company’s analytics division. There was a lot to consider. Between the motorbike traffic, lack of a subway (coming in 2020!) and inexorable heat, Saigon life is not without its challenges. But the food, coffee, nightlife and people I met in 3 days convinced me I could adapt to love the city.

Saturday morning I flew into Hong Kong. I had told myself that I shouldn’t move back to Hong Kong, that it wouldn’t be good for my career. But as my taxi zoomed down Gloucester Road past the pretty buildings, I thought to myself, I could carve a good life for myself here again. A few hours later I was in my party jacket hidden on the top floor of the Winery in Sai Ying Pun. It had been weird keeping my arrival a secret. As I heard the voices of my friends entering from below, I so desperately wanted to burst out and shout my presence. I managed to contain myself until Donna Doubet announced a surprise guest from America, and I flew down the stairs, threw some authentic New Orleans beads into the crowd, and awkwardly raised my arms, a little uncomfortable with the sudden attention. This was only my second time back in Hong Kong since I left, and I couldn’t imagine a better rewelcome party.

The next several days consisted of as many as 7 appointments a day, catching up with friends and family. I also got some intel on the work environment, and while the tech scene is growing rapidly, I confirmed my suspicions that interesting and high paying data science jobs don’t exist in Hong Kong yet. 
Doug popping his cherry blossoms
Doug’s party in Tokyo was a Hanami 見花, meaning cherry blossom viewing party, because of course Japanese has a word for that. We had rented out space at Yoyogi Park, roped off and carpeted to enforce a shoes-off policy. Organizer Niji had arranged for a small arsenal of whiskey, champagne and a buffet of sandwiches. To meet the formal dress code, I paired the New Orleans party jacket with grey suitpants. Even at this party, I met a programmer who tried to recruit me. I realized that having a skillset unbound by geographies or business domains could be a blessing and a curse - too many options means you have to restrict yourself to stay sane. I decided not to pursue working in Japan. 

Taking place a month past his 40th birthday, the Hanami was really a farewell party for Doug, as he had just accepted a promotion that would relocate him to Seattle. We spent the entire weekend bemoaning the challenges of managing an American career while smitten with Asia. Discussions with him and Austin taught me that for us, location can be more important than job, and wanting to learn a language is a legit factor in deciding location.

In Beijing I was fortunate to get connected with good tech people. My friend Joohee had moved from Hong Kong, and it was only face-to-face when I learned she now worked in Chinese tech venture capitalism. She connected me to the CEO of an AI startup trying to develop the flying car, and through another friend I met the former head of data science at Mobike. I learned about the speed of China’s 9/9/6 tech culture, the role of WeChat in everything, the way government-led directives influence entrepreneurs, and the sheer abundance of data available. Joohee evangelized her bullish views on China, and I was reminded how much I missed the uniqueness of Beijing life when I found myself telling my life story to an attractive group of film producers in Mandarin. I seriously wondered if I should focus on returning to China. However, barriers included the vast amount of competing Chinese programmers and the increasingly domestic nature of China’s tech scene that render multilingual people like me no longer highly valued. And this is before getting to all the moral and logistical complexities enforced by the Chinese government. 

By the time I got to London, I was exhausted. I met with friends there in interesting companies, and tried the city on for size. In my most productive conversation, I talked with a former coworker and ultimate teammate about how moving to Vietnam might mean missing weddings and/or an ultimate tournament in Amsterdam that I’d been invited to. “Oh, you have to go to Windmill.”  And so I did. I returned from my round the world trip with lots of renewed friendships, and lots of discussions comparing the social joys of living in Asia, the family warmth of staying in the east coast, and the asymmetric way America treats international experience. While Asia would always respect my US work experience, the converse was not necessarily true. I decided I needed to at least explore interesting jobs in the US to compare with the offer in Ho Chi Minh City.

-----

I was full of energy the first month back in Boston. I pursued all those things I never have the energy to do when working. I read voraciously, studied languages, attacked the gym, played ultimate and went through a TensorFlow tutorial.

The second month was harsh. The job interview process plodded along frustratingly slowly, and all my hard work towards self-improvement was largely irrelevant to the interviewers. It became difficult to sustain such intensity, and the uncertainty slowly ground me down. Not knowing where I’d live or what sort of income to expect made it difficult to plan things, date or try new activities. 

The Vietnam offer was still outstanding, while two local options were in play. One was a tech startup where I had wanted to work back in 2016 that was now recruiting me. They had given me a dataset assessment back then, and I laughed when they sent it again, virtually unchanged. With years of practice now under my belt, I did a way better job on the assignment. In the followup interview, a kid just out of college review the assignment with me. It was a little stunning to see 2018 as his graduation year, but he introduced to me a little trick transforming linear variables like date or time into cyclical variables, by taking the sine and cosine of them. The interview went well and they indicated they would bring me onsite. Then without explanation, they wished me luck and rescinded the onsite interview.

The second was a large tech firm where my friend had internally referred me as a product manager. I was excited to pivot away from straight technical work, which often strained my extroverted personality. That firm’s HR operated slowly, and weeks elapsed between followups. Finally in mid May, they brought me onsite for a marathon session of interviews. While the experience was largely positive, I reflected over the weekend and realized I needed to follow my heart to Vietnam. With that realization, I then booked my Europe trip for Windmill.

That following week my brother and sister-in-law visited, and I told my whole family that I was moving to Vietnam. They did not take the news well. They mainly believed that the low salary, distance from tech thought leadership and lack of any incredible valuation growth were wrong for me. Only my father, who had spent a decade working in Shanghai, considered the possible upside of being in a growing economy at the right time. The next day at breakfast, my brother asked me if I was happy at my last company. I had been, because we had been working towards real global impact, and that in the months since I hadn’t come close to any company that excited me like that. “Oh there was this company in New York that tried to recruit me last year. They’re using machine learning to solve city sustainability issues. Would that interest you?” “Uh, yeah, no shit that’d be cool.” “Damn, I should have remembered to bring this to you earlier.” “Yes, you should have."

I figured it was too late to apply as I had already interviewed onsite at the big firm. But my brother emailed the CEO, who responded extremely quickly, and the next day I spoke with the head of their urban analytics team. The conversation went shockingly well and I learned that this guy’s previous role was leading analytics for the city of New York. A Google search revealed him to be kind of a big deal, as well as a visible minority in the field. It’d be really cool to work for him. They sent along their dataset assessment and told me to take a week on it. I pulled an all-nighter and turned it back in a day and a half, producing some of the best modeling I had ever done including applying the cyclical transformation trick I’d learned just a couple weeks before.

Finally on Monday, a full two and a half weeks after I’d gone onsite, the big tech firm gave me an offer. It was at the level that I had wanted and legitimately thrilled me. It’s funny how much more interested I’d become in the role when the offer became tangible. Still, I made arrangements to speak with the Vietnam CEO and sent an email to the New York startup informing them of the offer. Again they got back to me right away, and I soon had a call scheduled with the CEO for Wednesday morning. The Vietnam CEO also asked to speak Wednesday morning, and I had a followup with the big tech firm for Wednesday afternoon. Wednesday evening I would fly out to Spain. It would be the most eventful Wednesday since the CiBO layoffs, and similarly, I couldn’t sleep at all the previous night. I had 3 separate timelines at my fingertips, with 3 very different cities and 3 very different roles.

The New York CEO informed me that they typically bring people onsite before offering roles, but he’d be willing to make an exception if I was committed. I replied that if they could meet my salary expectations, I’d also take an offer without coming onsite. He then said he’d have his people get back to me.

The Vietnam CEO and I had a good heart-to-heart chat, but he was not able to meet the salary expectation that I wanted. It was an enormous risk for both of us, and while I’m not exactly risk-averse, I realized that his startup probably wasn’t as ready as other places to get value of data science.

Finally, the Boston tech firm gave me their final offer and told me I had until Friday 5pm EST to accept it. After how long they took to get back to me, I was a bit resentful about the tight deadline they’d given me, but they had other candidates in queue.  I then proceeded to fly to Spain.

When I landed in Madrid on Thursday, I had emails from the New York startup. They wanted me to go on Google Hangouts with some more employees. Sigh. I wanted to vacation, but this was my future, so I said sure, how about 4pm EST/10pm Spanish time? Considering I’d done my final interview with CiBO in Tokyo, this wasn’t even unusual for me. I flew to Bilbao, met up with Antonio and his wife Raquel, grabbed a quick dinner and drink, then hustled back to get on Hangouts. The interview was full of challenging questions, but by now I’d done so many interviews I was almost on autopilot. In one of the last questions, they asked me how I approached a dataset. Tiredly, I asked back, did you see my assessment? Surprisingly, one of the interviewers excitedly responded, “yes, I thought it was awesome, it was so cool how you transformed those cyclical variables.” Fuck yeah, I thought. Finally I told them I had until 5pm tomorrow to respond to the tech firm.
¿Donde esta el email de Nueva York?
The next day we touristed around beautiful Bilbao. I tried to enjoy it as much as I could, but the whole day I was aware of the time in New York. 9am... they’re getting into work... lunchtime….no email yet. 11pm Spain was my deadline. By 9pm we posted at a bar with wifi. I tried to be a fun conversationalist but the anxiety was real. At 9:30, the New York startup sent me an email…all it said was “hang in there, we’ll get back to you within the hour.” By 10:30, they still hadn’t. 10:45, the inbox was still unchanged and I’d lost the ability to make conversation. At 10:50 Antonio lent me his phone and I called someone at the company. No response. Finally at 11:00, I sent an email to the big tech firm saying, “I accept!”

The burden was gone. I’d be a product manager in Boston. It’d be a good life. I approached the bartender and said, “Tres tragos de tequila por favor. Tengo un nuevo trabajo!”

I brought the three shots back to our table and prepared to do a toast to the new job. Glass in the air, I sneaked a peak at my phone and glimpsed one new email. “Wait hold on! We have an offer for you!” I put my glass down and sighed.

In the ensuing telephone call, I admonished the New York startup for being late. They quantified their offer, apologized for missing the deadline by 5 minutes, and asked me to consider rescinding the acceptance. That thought literally made my heart quiver - I hate going against my word. I sighed, told him I’d sleep on it, and to please send a formal offer via email. When I checked my email that night, there was a formal offer that was slightly larger than what he’d said on the phone - turns out accepting the other job was a good negotiating tactic - as well as a response from the big tech recruiter revealing her joy at my acceptance. 

I did sleep on it, sent the offer around to my family, and ended up choosing the New York startup - Urbint. The email to the big tech firm rescinding my acceptance was the hardest email I’ve ever had to write - I had to get my brother to draft it for me. I clicked send on the train to Paris, and now, two weddings and a painful move later, I’m in New York City.

The Calamity was longer than desired, but was an invaluable period reconnecting with friends. My lessons learned:
  • It is so valuable having a strong, diverse peer network to inspire and raise you
  • It's important to be patient
  • It's important to be bold
  • It's ok to prioritize location
  • Jobs are like buses. You wait around for ages and then they show up all at once
I've chosen to be patient, to put off my return to Asia for an exciting job opportunity. Hopefully I'll be here in Urbint and New York for a long while. If I'm in this position again in a year, I'll know I'm truly cursed. But if that happens, I'll tell big companies to pay me to work at their competitors.

Sunday, September 10, 2017

Emojis, elections, LKF and whatsapplied statistics

Whatsapp is the messaging medium of choice among the ultimate community of Hong Kong. Since December 2011, an ever-expanding group of players have coordinated practice, social activities, shared news and engaged in raucous discussions over a Whatsapp group. At times, this group has exploded like a phone ringing off its hook and been the source of much hilarity for the community. The group has gone through many names, but most commonly has been known as "Party in My Pants" (PIMP).

This analysis was started by my friend, teammate and former colleague Jak Lau. In June 2017 he emailed me this:

Party in my TUANsuit
STATISTICS
Started: 10 December 2011
Age: 5 years 7 months
Messages 39,700
Group name changes: 19
Messages written no.%
Doona32618.0
Sam24536.0
Jak22235.4
Neil21205.2
Mikey20975.1
Kim20415.0
Will13393.3
Tommy8052.0
Gio6551.6

This is the email exchange that followed:
Cal: What?!?!?! How did you get these??
Jak: I made them.
Cal: You downloaded the transcript and searched?
Jak: Yeah, just exported the chat into excel and used a few simple analysis tools. Mostly sort and filter.
I was gonna do more, but wouldn't be worth it.
Cal: Haha I might play around with it. I want to see which emojis we use (eggplant)
Jak: Have fun. I don't know if you can export the emojis.

And so I set forth to do some further analysis. Whatsapp's Export Chat feature is nifty (and a feature that separates it from many other messaging apps), allowing me to email my account's stored data of the phone to myself as one raw text file. You can include media as well (images and videos) but I figured that might be an overwhelming amount and didn't include the images. Since leaving Hong Kong in early 2016, I've both remained in this group chat and become a professional data scientist, learning many techniques that would help me work with this text file.

Python is a good tool for text analysis, especially when used through a web application interface like a Juypiter Notebook. The benefits of using a web interface is the text gets outputted in your browser, which means different language scripts and emojis, both of which are relevant here, will likely be supported. The Pandas package allows in one step for the ingestion of the .txt file and conversion into a useful dataframe. The raw data contains lots of activity from whatsapp, including when users entered or left the group and when they sent images. I was mainly just interested in messages. I knew I would eventually do analysis on the emojis though, and I wasn't sure how that would work in a Python Juypiter Notebook - there was a tutorial on emoji data science in R though, which happens to be my strongest programming language. I exported the raw data from Python and used R to analyze it.

So I imported the data into R and was looking at a 37,351 * 4 dimension dataframe. I had the timestamp of the message, the sender, the message text itself, and whether it was an image.

My first step was to look at a timeline - what was typical activity? What sort of spikes occurred and when? It's hard to plot a timeline without first grouping data into buckets, hence I grouped the continuous timestamps into the days of the message. Since I was in the US EST time zone when I emailed these messages to myself, the timestamps are also in that time zone - however the majority of the group is based in Hong Kong, and it made sense to add 12 hours to all those times first. Then I was able to calculate how many messages were sent on any given day in Hong Kong, and eventually create the following plot.

The labels that you see there were created semi-manually. After observing the timeline without the labels, I looked up the peak days and dug into the original data to see what people were talking about on those peak days. For most of these I could find some clear events that piqued interest in the group. Some of these were external events like the US Presidential Election and the Hong Kong Umbrella Revolution (perhaps the most sustained spike) and some were random internal events like the time we decided to play a game where people typed entirely in emojis and others guessed what movie they referred to. A few of the spikes didn't really correspond to anything more than a Saturday night. The external events were mostly major news events, live sporting events including ultimate events. I also felt like the chart was lacking colors and struggled to think of what other variable I could use to color the dots, and decided to just use day of the week. It isn't really adding any more insights to the chart, but it makes the presentation better.

This plot also shows several periods of no activity which I can trace to periods where I lost my phone and had to restore an earlier backup. Whatsapp stores data on the cloud of course but access to this data on any given device is local. Whenever I lost my phone and had to reset, weeks or months of messages were lost. This explains why my total message numbers are lower than Jak's, despite being done several months later. I also chose to break down the plot into each year, called facetting in ggplot, and keep each year to its own scale - the 469 behemoth during the first emoji movie game ruins the scale and makes it hard to see other outliers. This is where ggplot really shines - facetting is awful to do in many other plot packages, and without ggplot I would probably just create 6 plots and piece them together in a photo editor.

Next was to repeat Jak's work and look at the most common texters. The group has expanded greatly over the years. Once limited by the app itself to 25 users, it has now grown to 91 users, and I found unique messages by over 100 users including users who have left the group. After aggregating the sender count, I made the following bar chart including all texters who had sent over 100 messages. In R it was also easy to add in a bit of extra information by coloring the message quantities by year (which I do with sequential shades of green, making it clear that darker means more recent).

So we do see that Donna is far and away the most active user historically, followed by Sam, Mike Ying, Kim, Jak, Neil, me and Tuan. The group's expansion is also visible here, with users Wanda and Jason noticeable for being high volume users with messages only in 2016 and 2017. On the flip side there are users whose volume dropped off over the years, including people like Nickie Wong and Chris Harrison who left Hong Kong.

What was also fun was searching for specific words, and then redoing the barplot for just messages with those words. As one of the main functions of this group was to organize social activity among people in Hong Kong, several places in Hong Kong appear in hundreds of messages over the years. Chief among these is "LKF", short for Lan Kwai Fong, one of the best party areas in Hong Kong and in all seriousness, the world. LKF appeared in messages 102 times, led far and away by Tuan Phan.
Ok I'm #2, but Tuan has me beat by a mile. Along each bar, I included a randomly sampled message by the respective person using LKF, and it so happens that Ruth Chen's message is "Tuan's always in LKF."
Just as an added side bonus, I wanted to see how deep into the socializing these texts typically occurred. It took a bunch of manipulation (I had to extract  the time portions of these texts, then set them all to the same arbitrary day) before I was able to graph the frequency of these texts over the course of the day

Hmm, it would appear that texts referring to the party place in Hong Kong within the group "Party in My Pants" really take off between 6pm and 1am. Whodathunk it?

You can repeat the first LKF graph with any other word, or regular expression. I'll do one more, and be careful if you're reading this at work, because our group is not a PG13 group.
Interesting, Tuan also has a commanding lead in this category, and his randomly sampled "sex" sentence even includes "lkf." Even if you the reader are not familiar with any of the people mentioned here, you may have an inkling of why this groupchat is now named "Party in my Tuansuit."

Ok at this point, the most data science heavy thing I've done is sample a random sentence and plot it in that graph. Surely this is not what I'm paid to do (you'd be surprised). But let's actually apply some text mining to this wonderful data set.  I first do some quick preprocessing steps, reducing everything to lower case and getting rid of pesky punctuation. Using the R package "tm", I also eliminate a healthy group of English language stopwords (generic words like "an", "me", "who" etc which don't really provide any insight), and created a corpus and dictionary. Here dictionary means that the program creates a vector to store words. It iterates along each word of each message and every time it comes across a word, if it hasn't seen it before, it adds a new element to the vector and assigns it the value 1. If it has seen the word before, it finds the index corresponding to that word and increases its value by 1. The program will separately keep a vector of the words themselves so that we can match them to the word count later. This step tells me that we have 17,608 unique words. Considering we have over 37k texts and most texts have multiple words, I was surprised that the unique words was so low. As it turns out, we repeat words a lot. From this step I can see that we've said "happy" 1509 times and "birthday" 1215 times, the 1st and 3rd most used words respectively.

So I want to calculate the average frequency usage of each word, and the average frequency for each user. Key to this step is the document term matrix. The document term matrix is essentially a collection of all those word vectors but corresponds to each document, which in this case is an individual text. Each vector must have as many indices as there are unique words, so each vector is 17,608 elements long. Since there are 37,351 texts, we are looking at a 37,351 * 17,608 matrix! That matrix takes up at least 5 gb of ram on my computer. I say at least because my workstation would crash before it finished creating that matrix.

Luckily computer scientists have figured ways around this - a sparse matrix. Nearly all the elements in the matrix are 0 - no text has anywhere close to 17k unique words.  A sparse matrix only stores the non-zero elements. It is a little bit harder to do operations with this, but you still can, and it saves a lot of storage. The sparse matrix for this whatsapp group was only 5.3 mb. I combined this matrix with a vector containing the senders of each chat, and iterated through for each unique sender to find each person's total word vocabulary, or individual dictionary. These individual frequencies could be compared to the overall frequency in a couple ways. We could look at the % difference in values, finding cases where someone used a word 1/100th of the time and overall it was used 1/10000th of the time. However for words that were only used a couple times in total, this % value would be wildly distorted. So I removed from the consideration all words which were only used once overall, and created a weighted equation where the raw difference in values was also taken into consideration. The weightings I used here were arbitrary, but I tried a couple variations until I got words that seemed "interesting."

After successfully iterating through each sender (and not crashing my computer), I saved the 10 most "distinctive" words for each sender, and graphed these words for a bunch of people

Awesome. There are a lot of interesting words in here, which I'll get to in a bit. but first, what are the u0001---- things? Most of these are unicode for emojis - a couple of them are Chinese characters. And there are lots of emojis, to the extent that this graph is really more distracting than useful until those unicode sequences are converted into weird smiles. And thus I broke out the emoji data science tutorial, written by the affable Hamdan Azhar, who has actually founded a company around emoji analysis.

Turns out emojis are really complicated. The steps involve in making that graph pretty were extensive - I spent a couple weeks of free time on it. Hamdan's strategy is to create a dictionary mapping each unicode id to a name of the emoji, and downloading a bunch of emoji .png images with the same name. His tutorial links to a dictionary and set of images, unfortunately the dictionary I used did not contain unicodes and I had to find another one online. This one for some reason named some emojis differently. Aggravatingly differently. There isn't exactly one emoji regulatory body (or 👮👉 for short). For example, my dictionary had "grinning face with sweat" and my images used "smiling face with open mouth and cold sweat". Also, whatever regulatory body there is keeps adding new emoji and the dictionary and set of images were not up to date. So I expanded the repertoire as I came across new items. A new frustration came with the new png files, some of which create an error when I tried to render them. Turns out I needed to download windows specific emoji, some of which look quite different from browser, android or Apple versions. Eventually with enough "manual" work, I was able to redo the plot by removing all the text that matched with emojis, then one by one rendering an image of the matching emoji in their place. The cleaned up result is below:

How sweet is that? Some users almost exclusively communicate in emoji (Kingi, Cat MK, Rie). Mike Ying's first emoji definitely rang a bell, and of course the eggplant appeared on Neil's most distinctive words list. Also notable here are Sam Axelrod's 7th most distinct word, Lincoln's 4th and 7th, Clay's 9th, Conor talking about football, my love for Tom Brady, and the fact that Donna, the group's most prolific user, apparently just texts various different types of laughs all the time. Note, this emoji chart was done a bit later and with slightly improved methodology from the previous non-emoji chart, hence not all words match up.

More users:

Of course Wilkie mentions Madonna, Jeremy uses aviation vocab, Quention has a baby named "Marni" and Jason talks about master's. Ed Lee says "sold" whenever you propose any social event. I'm not sure what it says that Kirk's most distinctive word is "harem", but he's used it twice and no one else has. And yes there still are some frequency issues here. While I removed words that were used once overall, some of the words that show up here were only used once by the user and twice or thrice overall. Is a word really distinctive of a person if he/she has only used it once? Perhaps my weighting equation needs some reworking, but there was always going to be some issues, especially with users who haven't sent that many texts.

But Cal! There are still emoji unicodes in here! Yes there are. Like I said, emojis are really aggravatingly complicated. The basic emojis are all one unicode to one emoji - however they just keep expanding it. You know how the face emojis now have adjustable skin tones? That is a combination of two unicodes - the original face unicode and an additional one signifying the skin tone. All the flag emojis? They are a combination of two unicodes. And actually, England, Wales and Scotland are all considered subdivisions of a national flag and are somehow represented by a combination of 7 unicodes. My current methodology breaks up every unicode combination into individual words of one unicode. I could redo the process grouping everything into bigrams, but it's not even guaranteed to solve this problem. It's a tricky one that might be best solved with more manual work. The ungraphed emojis in the graphs above are mainly Hong Kong/USA/UK/England flags as well the skin-tone signifier emoji. There are still some Chinese text that are left encoded - while I've worked with Chinese text before, for some reason I had trouble getting them to display on this file. 

I get the impression that many people find big data and data science very abstract and impersonal. The algorithms crunching massive data behind your targeted ads don't exactly inspire congeniality. But these techniques can be applied to anything, including more personal data. I've already written posts looking at my Facebook data and my travel locations, where data visualization really helped me understand my own past better. Going through this particularly dataset was especially fun - I was constantly reminded of hilarious exchanges from years ago with friends on the opposite side of the globe. Does this analysis add any business value? Nope, but I spend plenty of time doing analysis for data that will add value, and sometimes it's fun to just see exactly how many times Tuan drunkenly messaged the group.

P.S. If I can do this, Whatsapp (Facebook) is also probably doing this with your data.

Tuesday, January 5, 2016

Travels Visualized

If you've read my post on my Data visualization of Facebook data, you may have learned that I have this odd hobby of playing with personal data. However you are unlikely to know that I have kept a spreadsheet of where I have spent every night since 2008. But it's true, and it hasn't even been remotely onerous. I have noted the date in which i entered a city and the date in which I left. I've kept the detail mainly to the city level, so there is no record of meanderings throughout urban bedrooms, if that had theoretically happened. There is no record of day trip cities I've visited without spending the night, so sorry Nara, Krebs and Brookline. I also did not stress about whether I reached a city before or after midnight - with the exception of red eye flights, if I flew from Hong Kong on a Friday night and reached my room in Taipei 2am Saturday morning, I recorded myself as spending Friday night in Taipei.

I started this list in the margins of my notebook in a Geography lecture in University College Dublin in the fall of 2008, mainly because I was bored in class. It was far more interesting for me to think about the places I'd traveled to in that very epic year of 2008. Through memory I filled in most of the dates, and eventually I went through my gmail archives and filled in my whole year. It would not be possible for me to fill in 2007 and before because of lack of memory and records. But starting from that fall, I created a spreadsheet and it has survived 4 computers and is now safely on the cloud. I started this blog in 2008 and I feel that I became a very different person starting from that year so I find this data very fitting.

For me, trawling back through this data just brings me pure joy. I had not planned on ever analyzing the data as I am doing now - just scrolling through it had been enough. I'm not sure if anyone else would enjoy it so much going through their own life, and they certainly wouldn't find it so interesting going over mine. But these data visualizations directly take me back to trips I had forgotten.
In terms of the visualization nitty gritty, I had a lot of cleaning up to do. I hadn't even heard of R when I began these records, so I didn't really have a thought to how I should format the data or the scheme of the database in technical speak.
Once I fixed name inconsistencies and date formatting, I decided to focus on the chronological part of my data first. I found the ten cities I'd spent the most time in, then realized many cities were tied and expanded the list to top 15. Without worrying about a y-axis at all, I plotted the dates I'd been in these cities and assigned a nice color palette to them.
And immediately I was pleased. As familiar with my own life story as I am, I could see a lot of stories in those dots. I was at first surprised to see what cities made it. Civate and Osaka are on there on the strength of one trip each, which were both for World's ultimate tournaments. Ultimate tournament trips to Manila are also clearly regular, spaced out evenly starting in 2012. My move from DC, where trips to New York and Newton (my hometown outside Boston) to Hong Kong in late 2011 is also quite obvious. My lengthier stays in Dublin and Beijing which are well-documented in this blog are also visible. Irregular trips to Shanghai, Taipei, Shenzhen and Bangkok pop up. Lastly, I may never spend another night in Newton after we sold our family home, and instead nights in the city of Boston show up instead.

These cities appear low to high in order of appearance (starting from 2008), which is really quite arbitrary. I played around with making the order completely random. I actually liked that better, because in the original version, all the long stays are at the bottom and the top seems very bare, giving the image a sense of imbalance.  My friend points out that at this point I enter data art, because the randomization serves no functional purpose. Here is that graph with even more cities.

Now this has more little stories and might be too cluttered. I don't think I could possibly fit all of my stops in there. Because I go back to places multiple times, I don't know how to provide a sense of chronological order to the cities. There isn't a sense in either graph really of many sequential trips.

Skipping this thought, I considered graphing the latitudes on the y-axis. As the data points were all geographic this was the logical next step. If you didn't recognize the names of these cities, the graph loses a lot of meaning. However, adding longitude and latitude coordinates was tricky. I hadn't been inputting that data in along the way, but the Maps package in R comes with a nice database of world cities. The database comes with population and coordinate information for just about every city of more than 40,000 people in the world (as of 2006). However the naming of some of the cities is bizarre, full of colonial-era names (Rangoon, Bombay), out-of-vogue spelling choices (Cracow, Soul), and a full pinyin rendering of Macau and Hong Kong, with Hong Kong being split into the Sai Kung, Kowloon and Hong Kong Island (xigong, jiulong and xianggangdao). Some of my travels had taken me to places of less than 40,000 people. For the names that didn't match, I didn't know of any better solution that editing all the names individually. If there is a more common sense laden database out there, please let me know. For the cities not on the database, I googled their coordinates. I'm not quite a good enough programmer to build a scraper to do this automatically, but I aspire to get to that level soon. Here's the same graph re-run with latitudes on the y-axis. I added in the names by hand where I saw fit, otherwise they'd bleed over each other.
This graph doesn't do much for me. It's very cluttered and it's hard to organize time and space coordinates together in my mind. The points give the illusion of coordinates, but they're not, they're only one dimensionally coordinates. It is interesting to see that so many of the places I've been are on the same latitude (Beijing is nearly the same as New York), and places that actually are near each other (Hong Kong and Shenzhen, Newton and Boston) now appear so on the map. And it's also cool to see that these cities range as south Bangkok to as north as Dublin, a thought I'd never made on my own. Still, I think there is limited use to this visualization attempt.

Well with all the coordinates in hand, it was time to put them on a real map. I've learned that with maps it's easy to control the size and color overlaid points. R has some in-built plotting functions, which I'd been using for 5 years, and they're pretty great. But people serious about data visualization seem to use a lot of ggplot2 and ggmap, packages developed by the Kiwi statistician Hadley Wickham, who should probably be knighted. I decided it was high time to learn some ggmap. Here's my first attempt:

This map visualization is much more of a classic representation and was equally delightful to me. I particularly liked the lack of borders in this version. It requires more deductive efforts in identifying all the points. I decided to color the points by the year that I first visited that point. Blue represents 2008 and takes up a lot of real estate. The size was set to the number of days I spent in each city. Note that the sizes are not scaled linearly, and in fact I haven't figured out how to control this properly. I actually am rather alright with the outcome though, otherwise Hong Kong and Washington, DC would drown out other cities. I could see more stories here, such as two separate trips to Europe. I could see the "outlier" travels of 2010 that took me to India and Peru. I saw two points from 2008 in western US and was confused as to what they might represent until I remembered I went to Las Vegas and Los Angeles many months apart that year. This truly shows the power of visualization, the ability to convey so much more in so little space. I could also see a bug in my data in this black dot in western America. Investigating further, I'd inputted the coordinates for Littleton, Colorado instead of Littleton, New Hampshire by mistake. Some of the years were also incorrect.

My second go at the map fixed these bugs and added borders just for kicks. I don't really like these large borders, which seem to drown out and minimize my destinations. But for the most part, I was very happy with this map. All I needed was a legend. Turns out this was way harder than I realized, because I had been using ggmap incorrectly this whole time. Hours of debugging later, I finally converted my years into factors and got this to popup.

As a conclusion, these Travels of Cal are about to get a lot more interesting. After 4 years and 3 months at Arup in Hong Kong, my last day will be January 8, 2016. I will be adding some data points in Southeast Asia and building on these visualizations. There's a greater purpose here than documenting my own life, rest assured.


Saturday, November 21, 2015

Personal Data Science

When I graduated with a Master's in Statistics in the summer of 2011, I had never heard of the term data science. And most of the world hadn't really either, it only picked up as a term really in the year following. Now I can no longer just say that sort of sentence. I have to show some data 

In these intervening years, I've been working in an engineering firm learning a lot about how buildings work, what sort of mechanical systems use power, how to pick a piece of glass with the right reflectivity, light and heat tranmissivity, how to model wind flow and all sorts of applied knowledge I never conceived of as a student. But I haven't been getting in on this data science action in my day job. 

But I did study statistics and I like to play with data. While trying to turn my academic knowledge into something with real world applicability, I've realized how off-base a lot of my education was. Let's start with undergraduate, where I took courses such as Abstract Algebra, Galois Theory, and Complex Analysis. I haven't even come close to using anything there at all. Even with multivariable calculus, one of the foundations of a math major, I can't remember Green's theorem and don't particularly care to. The statistics portion of our learning was thus certainly more applied. The regression course I took junior year has proved invaluable, and learning to code and model in graduate school has been great. The most valuable lesson I probably learned was to not overfit the data when modelling, to make sure your fundament processes are right rather than your results. However the fundamental processes of graduate statistics are flawed in this modern world - courses teach more like history courses. Painstaking attention is made to how a theorem was discovered and proved. Professors are convinced these steps are crucial. Sure, I think there is something to be said for understanding the theory behind a model or function, but goodness gracious we never use those proofs again. And we spend very, very little time working with real world data and never any of the large datasets that have become so common. There's some balance between blindly learning how to use a tool and understanding the tool's entire backstory and manufacturing process. 

And there's a ton of free data out there, but there's also my own data. There's stuff that's automatically kept track for me, like bank transactions, cell phone data etc. I decided to play with my own Facebook data. Due to Facebook's stringent API, you can't really access their stuff by a scraping algorithm. So I took my own data myself. Copy and paste. Went through all my statuses and took down the time of posting, the number of likes, comments, and then whether the status was a joke, pun, announcement, topical, language-related, a link, a check-in, a holiday etc. I know, I'm ridiculous. But I really wanted to practice and not lose out on the value of my education. While at the Census Bureau, I read a lot of statistics papers that I no longer recall, but I also attended a very popular talk by Dr. Nathan Yau, at the time a Statistics PhD student who had just published a book on visualizing data. He lectured on the value and techniques behind awesome data visualizations, and I was hooked. I bought his book and follow his blog (flowingdata.com). I still haven't come up with any graphics worthy of his blog, but I tried here. I started with a simple plot of the post likes vs time and colored them in differently based on some of the metrics I recorded. 

The data is actually quite fun to play with. There's a few motivating variables to analyze. For starters, I want to see how often I pun, and whether these tend to be the most popular posts. It turns out they're not! The graph to the left actually graphs 6 variables. There's time on the x axis and # of likes on the y axis. Blue dots are puns. The size of the dot indicates the number of shares the post has had (most have 0 and a few have 1), and actually the color of the dots are different for posts judged to be "topical." Posts that are squares are picture posts. However, I feel pretty strongly now that most humans can really only grasp 4 variables. Yeah if you stare long enough you can try to understand them all, but after 4 the mind really has to work. I redid the graph with a log y-axis, which I think looks a bit better but doesn't make the popular posts look as impressive. For the record, puns make up 18.6% of the last 3 year's posts. 

I also used the wordcloud package to pick some of my most used words. This is a good package and once I downloaded it, I really didn't have to do much. The results are quite pleasing and cool. Note that <97> represents some Chinese character - I can't get Chinese character display working on my R.

Well those days are pretty even, but Sundays are fun days aren't they. Sunday's have the highest mean likes, but this graph shows that Thursdays have the highest median. The results aren't very drastic though, and just looking at the sample variances one can see that the differences might not stand up to a test of robustness. Add in the fact that these times are all Hong Kong based but not all posts were, and I wouldn't publish an academic paper advising Thursday posts. As a note I'm a big fan of boxplots, but not everyone learns how to read them, and it seems like they might be a legacy of older statistics that'll seem too clunky in this new age.


If you notice though, none of what I've done has involved any sort of fancy modeling. Maybe my experience has been limited, but it seems most of the value of data science is relatively simple. Most often at work I'm asked "what's the average energy use for a tall office tower?" All I have to do is type in a simple query, but this is a service simply not available before we had the database. It doesn't involve any explanation to math illiterates, or model validation. With my Facebook example, just having the database in the first place setup to allow these useful queries is the main step. This is primarily why the data science game is shifted towards computer scientists right now. The market demand is to get data in all the right places rather than advanced statistical modeling, so you need programmers to scrape data or design apps or programs that continually feed in usable data and store it. You'll need a few Statistics PhD's scattered around to come up with the original algorithms, but everyone else just has to learn how they work.

Here's one more graph with comments and likes together:









But anyway, I want to just post my own favorite statuses. These are not the ones with the most statistical properties, just my own faves:

  • My mom asks me if I want a home sound system for my birthday - I tell her thanks but I'm not the stereo-type.
  • I'm being asked at work to write an "inception report." I'm not sure what that means, but I hope it doesn't involve a report within a report.
  • England is to football as Iraq is to civilization. Sure they might have invented it, but you could argue other countries are doing it better now.
  • The Hong Kong Football Association reportedly paid about $30million HKD to get the Argentine National team to come play in tonight's friendly against the 164 ranked Hong Kong team. I guess this is the second time this month that the government has paid a group of men to beat up on Hong Kongers."
  • Once upon a time, Georgetown and Syracuse had a rivalry. Georgetown won. The end.
  • Sometimes I do my own crossword puzzles, which I no longer remember. And then when I solve a particularly clever clue, I feel pumped that I solved the clue and more pumped that I wrote it. #lowselfesteem
  • Reading from a Kindle is not helping my shelf esteem.
  • China doesn't have a Mount Rushmore, it just has a Mao-nt.
  • Sometimes my phone says "Call Failed" and I think it says "Cal Failed" and I'm like come on, I really don't need you to rub it in.