flink-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "ASF GitHub Bot (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (FLINK-1901) Create sample operator for Dataset
Date Fri, 31 Jul 2015 16:15:04 GMT

    [ https://issues.apache.org/jira/browse/FLINK-1901?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14649419#comment-14649419

ASF GitHub Bot commented on FLINK-1901:

Github user tillrohrmann commented on a diff in the pull request:

    --- Diff: pom.xml ---
    @@ -224,6 +224,12 @@ under the License.
    +			<dependency>
    +				<groupId>org.apache.commons</groupId>
    +				<artifactId>commons-math3</artifactId>
    +				<version>3.5</version>
    +			</dependency>
    --- End diff --
    For that we have to add an entry in the `flink-dist/NOTICE` and `flink-dist/LICENSE` files.
But I can do that when merging the PR.

> Create sample operator for Dataset
> ----------------------------------
>                 Key: FLINK-1901
>                 URL: https://issues.apache.org/jira/browse/FLINK-1901
>             Project: Flink
>          Issue Type: Improvement
>          Components: Core
>            Reporter: Theodore Vasiloudis
>            Assignee: Chengxiang Li
> In order to be able to implement Stochastic Gradient Descent and a number of other machine
learning algorithms we need to have a way to take a random sample from a Dataset.
> We need to be able to sample with or without replacement from the Dataset, choose the
relative size of the sample, and set a seed for reproducibility.

This message was sent by Atlassian JIRA

View raw message