Skip to main content

Posts

Summer - Week 10

This week, I am working to get the dynamic time warping functionality into my program. The process of doing so includes re-processing the features to include the time series, putting each series back together when we construct sequences, and then performing the DTW to generate a number that will be used to compute the kNN of each sequence which can then be used for predictions with the models. The processing time of these activities has gone up significantly since we have been using five different metrics with each of the F phase datasets. I am returning to school next week, and once I've completed the DTW processing all that will remain before we put together our second paper (The date for the reach journal we would like to submit it to is October 1), I am hoping I will have time to look again into the Agglomerative Hierarchical Clustering concept, which I did not successfully complete when we explored it earlier in the summer and then changed focus to the paper. We heard back...

Week 23

This week, I completed my calculations for DAR and Longest sleep bout based on 'day' being the 'lights on' period and 'night' being 'lights off'. Based on our reading of other sleep studies, it seems that most sleep research focuses on 'lights off' periods as being of interest. One of the most prevalent metrics in the papers I've seen regarding abnormal sleep patterns is Sleep Onset Latency, which is the amount of time it takes for a subject to fall asleep after lights have been turned off. Since this is also something of a set time at St. Luke's (although we cannot be absolutely certain that the patient will attempt to sleep as soon as the lights turn off), it appears as of now to be the most relevant way to split the 24-hour period. After my initial research on sleep onset latency, I had hoped to calculate an 'individualized sleep time' similar to the way we calculated features based on individualized wake time. Aft...

Week 22

This week, I finished up the graphs of Coefficient of Variability that Dr. Weeks had asked for and worked on calculating sleep onset latency and then the change in onset latency from night to night. Once I have these results, I want to apply this comparison to the different features that we've previously calculated (DAR, Longest Sleep Bout, etc.). I also am looking at the average stay of patients. Going forward, we're going to work on classifying sleep features as positive or negative with the help of Dr. Weeks and Dr. Skornyakov and see what new challenges this approach presents.

Week 21

Dr. Weeks and Dr. Skornyakov have recruited three grad students to help with data collection. Alexa and I are still helping out, and intend to continue through the end of February. Currently, we have a single patient in the study. I am struggling with some technical difficulties and was not able to produce graphical results of my coefficient of variation calculations, but hope to do so this week after downloading a Windows virtual machine or spending some time in the lab. Looking forward, I hope to get the CV calculation squared away and sent to Dr. Weeks. After spending time away from my exploration of individualized sleep time, I think it would be interesting to try adding an individualized sleep time as well. I will be consulting Dr. Sprint about this subject to see if it is necessary and how I could determine whether or not it was likely a more accurate version of sleep time than a standardized Lights Out time.

Week 20

I am wrapping up my individualized sleep calculator and will begin my individual sleep stats calculator. This proved to be difficult with the way I had initially set up my features calculator since I was using the Day/Night Lights On/Lights Off times determined by Dr. Skornyakov as constant values. In order to do this properly, I am going to have to essentially redo my calculator to better support a variable wake time. This is probably something I should have done from the start, and going forward I hope to not subject myself to the same tedium by minimizing the number of constants I use. Dr. Weeks has suggested that the most effective way for us to determine "good sleep " from "bad sleep", we should develop some baseline value that we can use to determine whether or not the patient is doing better or worse relative to only their own behaviour. To recognize how a value can be compared to the "usual" behaviour, we will say that our baseline period will...

Week 19

The Guide to Actigraphy data collection proved to be very helpful! We'd come across many articles comparing Actigraphy to Polysomnography, but few that discuss research using solely watch-based data collection. The study by Dr. Weeks and Dr. Skornyakov has not relied on Actigraphy logs, as is the general recommendation, because their subjects are not going to be consistently capable of filling them out with accurate information. The issue then will be determining whether our Day/Night times are appropriate for the individual patient. I think that Dr. Skornyakov's individualized sleep idea will see to this issue fairly well. An additional factor I have seen over and over again during our literature review is sleep onset latency. This is what is used in sleep labs to diagnose things like narcolepsy and insomnia. It is a measurement of how long it takes a subject to fall asleep once they lie down in complete darkness. Because we are dealing with patients who spend more time i...

Week 18

Over Christmas break, I am going to focus on doing more literature review. Dr. Skornyakov has provided us with some interesting articles on sleep studies of subjects suffering from disordered sleep. One, in particular, that seems of interest to me is a brief manual on what collecting and interpreting Actigraphy data should (ideally) look like in a medical setting. The others focus on topics not exactly the same as ours, but all involve some form of evaluating sleep. Unfortunately, two of them rely on cognitive testing (good brain function should indicate good sleep). However, it could be interesting to look at what factors of collected data they associated with high brain activity. In addition to reading, I will continue to work on the Individualized wake time calculator.

Week 17

At the request of Dr. Skornyakov, Alexa, Dr. Sprint and I intend to calculate individualized wake time factors for each day. This will be determined by the patient's sleep/wake status between six and seven in the morning. The idea behind this is that previously if a patient had woken before our "Day" period began (7:00AM) their activity would be recorded in the previous day's sleep period. Dr. Skornyakov has provided clear instructions on how to calculate this time. I plan to start by adding an extra level of processing to the output files such that we will have an additional column all for this new factor. Then, I'll modify my feature calculator so that it accesses the features from the newly determined time period.

Week 16

After last week, I had successfuly put together my script to calculate the length of the longest bout of sleep for each patient for each night during the study, but I was not displaying my output in a very meaningful way. This week, I was able to take the numbers I was getting and translate them into something that was somewhat insightful. The first obstacle I discussed in my last post was the lack of consistent length in the patient's stay times. To overcome this, I came up with a way to standardize the stay of each patient's data to seven days. I used integer division to divide the stay length by seven to get a new 'day length'. I then found how many days remained that did not fit evenly into the 'day length' and distributed the extraneous days so that in many cases, the 'day length' for the first few days in a patients stay were a day longer than that in the last few days. By averaging the longest bout over each 'day', I was able to produce ...

Week 15

This week, I began working on calculating a new data feature suggested to us by Dr. Skornyakov and beginning to investigate the significance of the longest bout of sleep in a night. The latter became of interest to me after reading the same paper provided by Dr. Sprint several weeks ago that introduced the concepts of the DAR (daytime activity ratio). The study calculated an additional feature that aimed to capture how fractured a patients sleep was. While my partner and I were unable to reproduce the feature's calculation on our data due to some ambiguity in its description, it still got us thinking about paying attention to the degree of consistency of sleep. I am going to start by locating the longest bout of sleep during a night and determining its length. Once I have this working well, there are several other angles that I think would be interesting to explore: start times of longest bouts length of other bouts considered relative to the longest length of longest bout r...

Week 14

In the past week, I managed to finish up my longest-bout calculators for both the longest sleep and longest wake periods. The results for longest sleep bouts were significantly lower than I would have predicted, so I spent some additional time trying to verify my results. According to my calculator, many patients had an average longest bout that was less than a hundred minutes - that is to say, throughout the night they never slept for more than two hours at once! While we expect somewhat irregular sleep patterns from our subjects due to their traumatic brain injuries or strokes, this number seems extreme. However, upon checking the data for a single patient in excel, it did not appear to be incorrect based on our current criteria for sleep/wake. In the next week, I'm going to check on an additional subject's data - one with the lowest average longest sleep bout - to ensure again that I'm not making a mistake. Then, I am going to work on standardizing and modelling the da...

Week 13

This week, I wrapped up my daytime activity ratio calculator and spent time reviewing literature that may shed light on how we can determine something like good sleep from our data. We have decided that if we cannot find a data-driven way to label sleep good or bad, we will have to refer to the self-reported sleep questionnaires that are given to subjects every third day of their time in the study. Our primary reference would be the Karolinska Sleepiness Scale, which asks the subject to select a face in a lineup of five faces that they believe best represents how sleepy they feel at the moment. These questionnaires are provided at consistent times for each patient at a consistent rate. In the upcoming week, I am going to continue our literature review and the search for meaningful features.

Week 11

This week, I completed the all-subject feature calculator. I do not have a lot of new information to report on this front, as we are calculating all of the same features we were doing for one individual subject. The next step will be to start looking through these features and seeing what parts of the data might be useful to us going forward in our classification. Dr. Sprint has given Alexa and I a copy of a sleep study in which sleep quality was measured by something called the Daytime Activity Ratio. This calculated what proportion of the patient's activity occurred during the day. We both plan to incorporate this new feature into our calculations to see what sort of results we get.

Week 10

As mentioned last week, I spent some time this week getting up to speed on github and git command line tools. I definitely wouldn't consider myself an expert, and the concept of branching is still quite foreign to me, but I'm confident enough currently to work on Alexa and my repository without fear of deleting everything. We continue to not have patients enrolled in the study at St. Lukes. This upcoming week, I will continue my work on the all-subject statistic calculator and reading on sleep classification.

Week 9

Wrapping up last week, I have my stat-calculating script working well and am ready to move forward. My partner, Alexa, and I have decided that we should start to begin working collaboratively on the next step of the project involving multiple data frames. To do so, we will be using a github repository. I have limited experience with the git tools, so I took time in the past week to try and gain a better understanding of the functionality that will be available. Alexa and I have also decided that instead of taking a break from the project itself to begin our literature review, we want to have some measure of both each week. We'll start by taking a more in-depth look at many of the sources utilized in the proposal for our project and studies published by Doug Weeks and Gina Sprint that are related to our current work. We hope to follow source material from those papers to find more relevant literature going forward. Dr. Weeks discharged the two patients that were enrolled in the ...

Week 8

My goals for this week are to:     1. re-create the slicing functionality of the stats function that I lost last week     2. begin working on a program that will allow me to analyze multiple data frames at once The stat-calculator revamp is coming along smoothly, but the multi-frame analysis is still in the planning phase. I intend to re-write my Automated Sleep/Wake analyzer to clean up any unnecessary or overcomplicated features of the original. In addition to working on the programs, we've got two new patients enrolled in the light study so I will be spending a few of my mornings at St. Luke's. Looking forward, after I am able to put together a working multiple-frame analyzer, I'll work on the literature review portion of the project.

Week 7

There are no patients currently enrolled in the study at St. Lukes, so this week's work was entirely focused on coding and reading about slicing. This was my second week of work dealing with the summary statistics of the practice data for subjects K002 and K027. My original code was functional and produced a result that was just slightly different than my mentor's. However, I'd chosen quite a roundabout way of doing this. While I did manage to create a date-time index for my DataFrame, I failed to fully utilize the full potential of this set-up. Instead of passing a slice of the DataFrame through a function for each period, I was moving through the data line by line in order to determine the number of transitions, minutes of sleep and minutes of activity for the period. Clearly, the former is both more efficient and more modular than the latter. I planned to create a function slice_stats that would accept a slice of a DataFrame up to twenty-four hours in size, but sho...

Week 6

Once I managed to get the frame clean and my code working, I checked my results against Dr. Sprint’s to ensure their accuracy. There were a few minor tweaks to make, but the vast majority of my output matched hers. We compared our results for two different subjects to ensure that each case was tested. Now, I’m ready to start calculating the daily statistics for my output. Firstly, I need to change my DataFrame’s indices to DateTime objects so that I can reference each epoch by its time, not an arbitrary index.     My plan for collecting the statistic data is to create a new DataFrame with columns corresponding to each calculated value and rows representing each day during the period the watch was worn. In order to ensure that one “day” will capture an entire day and an entire night, our clock will begin at 7:00 am and end at 6:59 am the following real-day. The day is then further classified into “lights-out” and “lights-on” periods, which I will also be calculating summary st...

Week 5

I’ve successfully implemented the sleep-checking script on my test data, but I’m having a bit of trouble transitioning to the real thing. Before I can even run my code, I need to clean up the data and ensure that I get the correct header row while dropping all blank rows and N/A columns. This has been trickier than I expected, as I’m having a bit of trouble discerning the proper indexing method (.loc[], .iloc[], or []) in each situation. I am confident, however, that once I am able to set up the frame properly, my algorithms should work effectively.     I’m going to spend the next week working on cleaning the DataFrame and picking up any loose ends that remain. My next objective will be to begin calculating summary statistics for the data – minutes of activity in a day, minutes of sleep in a day, etc..

Week 4

 Now that I’m more confident using DataFrames, I’m beginning to practice using them in a small-scale version of the sort of analyses we’ll be using later in the project. My objective is to write code that will use a series of rules to check each minute of Actigraph data and determine whether or not the wearer was asleep or awake at that point in time.     Before I test my code on the actual data, I’ve created a test dataset and found the desired result by hand. This way, I can test each rule and ensure that it is working before I attempt to process the massive data set.

Week 3

    I didn’t return to St. Luke’s until this weekend. My cohort, Alexa, was able to join me in shadowing Sarah, the remaining research assistant. Dr. Weeks had consented a new patient during the week, and another had just been discharged. After going through the procedures again with Alexa, I am fairly confident in my ability to complete the research tasks on my own. I will be doing so later this week when I discharge the new patient.     My work on the online course material has continued at a steady pace. I have entered the section involving NumPy, SciPy, and Pandas. In the upcoming week, I hope to gain a more in-depth understanding of these.

Week 2

The third and final element in project preparation was one that I did not see coming initially. When Dr. Sprint initially discussed the potential research topic with me, she explained her background in the area of activity analysis and association with Dr. Doug Weeks, the Director of Clinical Research at St. Luke’s Rehabilitation Institute. Dr. Weeks was involved in several projects utilizing data from wearable Actigraph watches to study the recovery patterns of stroke and traumatic brain injury, patients. Dr. Weeks’ research assistants were two volunteer physical therapy students who were due to begin clinical rotations at the end of the summer. Thus, it made sense for Alexa and me to replace them. I met with one of the students, Ellie, on a Saturday morning. She explained how there were two patients in the study at current, but that they typically had anywhere from one to three.      Moving forward from this training, my goals will be to get comfortable with the da...

Week 1

I will summarize my goal regarding my preparation for our research project with the following objectives: 1.    to obtain a broad, working understanding of Python 2.    to learn how to best utilize Python libraries and tools specific to data analysis 3.    to become a helpful member of the student research team at St. Luke's Rehabilitation Institute, where our data collection will occur My Python background is very limited. Prior to my current year of school, I had only taken a single class that involved it. In that class, we focused solely on simple math calculations and basic graphing functions. In the two years since then, my knowledge of the language shrank down to almost nothing. My memories of it, however unspecific, are fond ones. To accomplish my first task, I started working through Learn Python the Hard Way, by Zed Shaw. The book was designed to be intelligible to anyone, regardless of previous coding experience, so there was quite a bit that...