Commit Graph
330 Commits
Author SHA1 Message Date
Scott Sanderson bddb453272 BUG: F.window_safe implies f.demean().window_safe. 2016-09-22 12:41:50 -04:00
Scott Sanderson ab9a5d7060 MAINT: Use randint instead of random_integers. 2016-09-20 17:12:09 -04:00
Scott Sanderson 46cf54b180 MAINT: Remove outdated compat code. 2016-09-20 17:12:07 -04:00
Scott Sanderson a9c02935c6 Revert "MAINT: Remove support for custom string Column missing values."
This reverts commit 1b1e842e2339d6d0ee40cdfe34dcd27b4e4a7c0c.
2016-09-20 17:12:07 -04:00
Scott Sanderson ed365dc5fe MAINT: Remove support for custom string Column missing values.
Pandas 0.18 deprecated passing "null-ish" values to pd.categorical.  The
expectation, instead, is that you use categorical's native support for
missing data, which means the user will always get NaN's for missing
entries of the categorical.

A follow-up to this change should probably drop support for custom
missing values entirely and to use LabelArray/categorical for integer
data.
2016-09-20 17:12:07 -04:00
Scott Sanderson 9aa866e434 MAINT: Use sort_values() instead of sort().
pd.DataFrame.sort() is deprecated.
2016-09-20 17:12:07 -04:00
Scott Sanderson 7c58bf1e5b STY: Flake8 and parameter rename. 2016-09-14 14:45:00 -04:00
Scott Sanderson a0e1b881aa MAINT: Move refcount management into TermGraph. 2016-09-14 11:16:40 -04:00
Scott Sanderson 6aa885dbeb PERF: Release unneeded pipeline terms.
Refcount pipeline terms during execution and release terms once they're
no longer needed.

This dramatically reduces memory usage on large pipelines.
2016-09-13 23:28:25 -04:00
Scott Sanderson 8b2446aec6 ENH: Dont allow length=1 regressions/correlations.
They're not meaningful, and they cause warnings from numpy.

Implemented in terms of a new preprocessor, `expect_bounded`, which
takes a tuple of `upper_bound` and `lower_bound`.
2016-09-02 12:49:09 -04:00
Richard Frank bc87ea4efb MAINT: Consolidate coercion to sqlite conn/eng 2016-08-24 13:24:07 -04:00
LotannaEzenwa c19d2855ad Added kwargs to ewma constructors
fixed typos
2016-08-22 13:14:23 -04:00
Scott Sanderson bdc72ec4c0 BUG: Fix broken graph visualizations. 2016-08-18 11:07:17 -04:00
Scott Sanderson c5a3dae267 BUG: Supply a module for Downsampled terms. 2016-08-17 19:48:33 -04:00
Scott Sanderson 115f055c83 MAINT: Clean up downsampling boilerplate.
Consolidate docs and mixin applications into one place.
2016-08-17 16:52:09 -04:00
Scott Sanderson 8fa51bdaab MAINT: Use numpy_utils.as_column in more places. 2016-08-17 16:52:09 -04:00
Scott Sanderson d917a64d45 ENH: Add non-windowed downsampling. 2016-08-17 16:52:09 -04:00
Scott Sanderson a9b5b2ac2f DOC: Docstring cleanups. 2016-08-17 16:52:09 -04:00
Scott Sanderson 5f686173f1 STY: Flake8 cleanup. 2016-08-17 16:52:09 -04:00
Scott Sanderson 91276c7274 ENH: Add support for downsampling.
Adds a new ``downsample`` method to all computable terms.  Computable
terms (Filters, Factors, and Classifiers) can be downsampled to yearly,
quarterly, monthly, or weekly frequency.

The result of ``term.downsample`` is a new term of the same
family (Filter/Factor/Classifier) as ``term``.  The downsampled term
computes by delegating to the original term; repeatedly calling its
``compute`` method with length-1 date ranges.

Downsampled terms take advantage of a new ``compute_extra_rows`` Term
method, which allows terms to dynamically request that additional extra
rows of themselves be computed based on the dates for which they're
being computed.  This ensures, for example, that a monthly-downsampled
term always computes at the start of a month, even when a
naively-calculated pipeline window would end in the middle of the month.
2016-08-17 16:52:09 -04:00
Scott Sanderson 1444a78330 MAINT: Refactor in prep for downsampled terms.
- Split out extra_rows handling into an `ExecutionPlan` subclass.
  `ExecutionPlan` now requires the dates and calendar against which a
  set of terms will be computed, and now defers to a term's
  `compute_extra_rows` method when deciding how many extra rows are
  required to compute for that term. This will allow downsampled terms
  to request enough extra rows to guarantee that we can maintain consistent
  calculation dates.

  As a consequence of the above, `TermGraph` now only deals with logical
  dependencies, not with metadata surrounding extra row calculations.
  This means that TermGraph can be used to generate dependency
  visualizations in interactive contexts where we don't yet have a
  calendar or start/end dates.

- Refactored test_{filter,factor,classifier} to use check_terms instead
  of run_graph.  This makes it easier to make changes to TermGraph,
  since the testing interface is now to simply provide a dict of terms.

- Refactored BasePipelineTestCase to use fixtures to create an asset
  finder.  This fixes a potential leak of the test's asset db, which was
  not being explicitly cleaned up.

- Refactored test_technical to use BasePipelineTestCase.

- Added a new special term, `InputDates()`, which can be used to request
  date labels for inputs.  Like `AssetExists`, `InputDates` is provided
  in the initial workspace by default.

- Added a default (failing) `_compute` method to `AssetExists` which
  provides a more useful error than AttributeError.
2016-08-17 16:52:09 -04:00
Scott Sanderson 62d69db7f6 MAINT: Remove empty inputs from BoundColumn.
They belong on LoadableTerm instead.
2016-08-17 16:52:09 -04:00
Scott Sanderson 765f9b6d57 MAINT: Improve/test errors for insufficient data. 2016-08-17 16:52:09 -04:00
Scott Sanderson 4c59857e1f DOC: Add a docstring for RecarrayField. 2016-08-17 16:52:09 -04:00
Scott Sanderson 14a95449ca DOC: Clarify how AssetExists() is special. 2016-08-17 16:52:09 -04:00
Scott Sanderson b6bacd2815 DOC: Fix typo in docstring. 2016-08-17 16:52:09 -04:00
dmichalowicz 1dad512184 BUG: zscores should be window safe 2016-08-08 18:07:34 -04:00
Gil Wassermann 483397e554 ENH: Added AtLeastN filter 2016-08-02 16:34:32 -04:00
dmichalowicz 97099a0e92 DOC: regression docstring typos 2016-08-02 11:14:41 -04:00
Scott Sanderson f13294de4e ENH: Rename StrictlyTrue to All and add Any().
Also, moved All() and Any() to `zipline.pipeline.filters.smoothing`.
2016-08-01 22:10:28 -04:00
Gil Wassermann 7623c0f6eb MAINT: .sum() behaviour 2016-08-01 13:48:14 -04:00
Gil Wassermann 73de8e6182 STY: style changes and strictly_true_filter 2016-08-01 11:16:02 -04:00
Gil Wassermann 694d9e952a ENH: added smoothing to zipline 2016-08-01 08:20:10 -04:00
Scott Sanderson 161771917e DOC: Mention groupby in top/bottom docs. 2016-07-26 02:57:35 -04:00
Scott Sanderson 49bb8264dc ENH: Finish adding groupby to rank/top/bottom.
- Added test coverage for grouped and masked top/bottom.

- Added test coverage for grouped rank on datetime factors.

- Fixed an issue where grouped rank would fail on datetime inputs
  because unary-negative isn't defined for datetimes.  We now instead
  directly invoke a function from rank.pyx that does the normalizations
  as neeeded.

- Fixed an issue where GroupedRowTransform assumed that it produced the
  same dtype as its input.  This isn't true for rank() of a
  datetime-dtype factor.  GroupedRowTransform now takes a required dtype
  parameter.

- Similarly, fixed an issue where GroupedRowTransform assumed that its
  missing_value was the same as its parent's, which isn't true for
  rank() of a datetime-dtype factor.  GroupedRowTransform now takes a
  required dtype parameter.

- Fixed an issue where Factor.demean() and Factor.zscore() weren't
  properly cached because their static_identity included a closure that
  was dynamically generated on each invocation.  They both now always
  use a function defined at module scope.
2016-07-26 02:57:35 -04:00
Andrey PortnoyandScott Sanderson 9e3404646e add groupby to rank, top, and bottom 2016-07-25 23:53:33 -04:00
Joe Jevnik 25474cf475 DOC: add default inputs and window length to TrueRange 2016-07-25 12:37:25 -04:00
ChrisPappalardoandJoe Jevnik 072ada812a STY: fixed flake8 failing test 2016-07-25 12:37:25 -04:00
ChrisPappalardoandJoe Jevnik 5888cf1657 ENH: add true range technical factor 2016-07-25 12:37:25 -04:00
Scott SandersonandGitHub 75b3dc6fe4 Merge pull request #1338 from quantopian/window-safe-filters
ENH: made filters window safe
2016-07-25 10:38:08 -04:00
Scott Sanderson be857ead0e DOC: Clarify default window-safety for Filters. 2016-07-24 21:17:16 -04:00
Scott SandersonandGitHub a424225dce Merge pull request #1345 from quantopian/notnull-filter
ENH: Add NotNullFilter.
2016-07-24 21:15:08 -04:00
Scott SandersonandMaya Tydykov 43957b0d09 ENH: Add NotNullFilter. 2016-07-22 15:32:46 -04:00
dmichalowicz f404538008 DOC: More pipeline docstring tweaks 2016-07-22 13:54:16 -04:00
Gil Wassermann 98be158c20 ENH: storing commits. test case added 2016-07-21 08:49:41 -04:00
Gil Wassermann d7b631617c ENH: made filters window safe 2016-07-20 17:10:08 -04:00
dmichalowicz 9cc5796b3e DOC: Pipeline docstring edits 2016-07-20 15:10:23 -04:00
dmichalowicz a8486c5f6e ENH: Factor-to-factor correlations/regressions 2016-07-19 11:16:55 -04:00
Jean Bredeche 5a0f840917 Clean up daily bar reader/writer to take advantage of new trading calendar. The reader
is backwards-compatible with the previous format.

In USEquityLoader, use dailyreader's trading_calendar.

This is backwards compatible and will fall back to the NYSE calendar if
the reader doesn’t have a calendar specified.
2016-07-15 15:13:57 -04:00
Joe JevnikandGitHub 835fab8ebd Merge pull request #1323 from quantopian/pmap-blaze-query
ENH: Adds the ability to run blaze queries concurrently
2016-07-14 18:40:57 -04:00