原文:
The data is related with direct marketing campaigns (phone calls) of a Portuguese banking institution. The classification goal is to predict if the client will subscribe a term deposit (variable y).

Data Set Information:
The data is related with direct marketing campaigns of a Portuguese banking institution. The marketing campaigns were based on phone calls. Often, more than one contact to the same client was required, in order to access if the product (bank term deposit) would be ('yes') or not ('no') subscribed.
There are four datasets:
1) bank-additional-full.csv with all examples (41188) and 20 inputs, ordered by date (from May 2008 to November 2010), very close to the data analyzed in [Moro et al., 2014]
2) bank-additional.csv with 10% of the examples (4119), randomly selected from 1), and 20 inputs.
3) bank-full.csv with all examples and 17 inputs, ordered by date (older version of this dataset with less inputs).
4) bank.csv with 10% of the examples and 17 inputs, randomly selected from 3 (older version of this dataset with less inputs).
The smallest datasets are provided to test more computationally demanding machine learning algorithms (e.g., SVM).
The classification goal is to predict if the client will subscribe (yes/no) a term deposit (variable y).
Attribute Information:
Input variables:
# bank client data:
1 - age (numeric)
2 - job : type of job (categorical: 'admin.','blue-collar','entrepreneur','housemaid','management','retired','self-employed','services','student','technician','unemployed','unknown')
3 - marital : marital status (categorical: 'divorced','married','single','unknown'; note: 'divorced' means divorced or widowed)
4 - education (categorical: 'basic.4y','basic.6y','basic.9y','high.school','illiterate','professional.course','university.degree','unknown')
5 - default: has credit in default? (categorical: 'no','yes','unknown')
6 - housing: has housing loan? (categorical: 'no','yes','unknown')
7 - loan: has personal loan? (categorical: 'no','yes','unknown')
# related with the last contact of the current campaign:
8 - contact: contact communication type (categorical: 'cellular','telephone')
9 - month: last contact month of year (categorical: 'jan', 'feb', 'mar', ..., 'nov', 'dec')
10 - day_of_week: last contact day of the week (categorical: 'mon','tue','wed','thu','fri')
11 - duration: last contact duration, in seconds (numeric). Important note: this attribute highly affects the output target (e.g., if duration=0 then y='no'). Yet, the duration is not known before a call is performed. Also, after the end of the call y is obviously known. Thus, this input should only be included for benchmark purposes and should be discarded if the intention is to have a realistic predictive model.
# other attributes:
12 - campaign: number of contacts performed during this campaign and for this client (numeric, includes last contact)
13 - pdays: number of days that passed by after the client was last contacted from a previous campaign (numeric; 999 means client was not previously contacted)
14 - previous: number of contacts performed before this campaign and for this client (numeric)
15 - poutcome: outcome of the previous marketing campaign (categorical: 'failure','nonexistent','success')
# social and economic context attributes
16 - emp.var.rate: employment variation rate - quarterly indicator (numeric)
17 - cons.price.idx: consumer price index - monthly indicator (numeric)
18 - cons.conf.idx: consumer confidence index - monthly indicator (numeric)
19 - euribor3m: euribor 3 month rate - daily indicator (numeric)
20 - nr.employed: number of employees - quarterly indicator (numeric)
Output variable (desired target):
21 - y - has the client subscribed a term deposit? (binary: 'yes','no')
译:
这些数据与葡萄牙一家银行机构的直接营销活动(电话)有关。分类目标是预测客户是否会认购定期存款(变量y)。

数据集信息:
该数据与葡萄牙一家银行机构的直销活动有关。营销活动以电话为基础。通常,需要与同一客户进行多次联系,以了解产品(银行定期存款)是否认购(“是”)。
有四个数据集:
1)bank-additional-full.csv,包括所有示例(41188)和20个输入,按日期排序(2008年5月至2010年11月),与[Moro等人,2014]中分析的数据非常接近。
2) bank-additional.csv,10%的示例(4119),从1)和20个输入中随机选择。
3)bank-full.csv,包含所有示例和17个输入,按日期排序(此数据集的旧版本,输入较少)。
4)bank.csv,包含10%的示例和17个输入,从3个数据集中随机选择(该数据集的较旧版本,输入较少)。
提供最小的数据集来测试计算要求更高的机器学习算法(例如SVM)。
分类目标是预测客户是否会认购(是/否)定期存款(变量y)。
属性信息:
输入变量:
#银行客户数据:
1-年龄(数字)
2-工作:工作类型(分类为:'管理员'、'蓝领'、'企业家'、'女佣'、'管理'、'退休'、'自营职业'、'服务'、'学生'、'技术员'、'失业'、'未知')
3-婚姻:婚姻状况(分类:“离婚”、“已婚”、“单身”、“未知”;注:“离婚”指离婚或丧偶)
4-教育(分类为:“基础4y”、“基础6y”、“基础9y”、“高中”、“文盲”、“专业课程”、“大学学位”、“未知”)
5-违约:是否存在违约信用?(分类:“否”、“是”、“未知”)
6-住房:有住房贷款吗?(分类:“否”、“是”、“未知”)
7-贷款:有个人贷款吗?(分类:“否”、“是”、“未知”)
#与当前活动的最后联系人相关:
8-联系人:联系人通信类型(分类:“蜂窝”和“电话”)
9个月:一年中的最后一个接触月(分类为:一月、二月、三月、十一月、十二月)
10-day_of_week:一周中的最后一天(分类为“周一”、“周二”、“周三”、“周四”、“周五”)
11-持续时间:最后一次接触持续时间,以秒为单位(数字)。重要提示:此属性对输出目标有很大影响(例如,如果duration=0,则y='no')。然而,在执行呼叫之前,持续时间是未知的。而且,在呼叫结束后,y显然是已知的。因此,仅应出于基准目的纳入该输入,如果目的是建立现实的预测模型,则应放弃该输入。
#其他属性:
12-活动:此活动期间和此客户端执行的联系人数(数字,包括最后一个联系人)
13-pdays:上次联系客户后经过的天数(数字;999表示之前未联系客户)
14-上一个:此活动之前为该客户执行的联系人数(数字)
15-poutcome:上一次营销活动的结果(分类:“失败”、“不存在”、“成功”)
#社会和经济背景属性
16-环境风险率:就业变化率-季度指标(数字)
17-缺点。价格。idx:消费者价格指数-月度指标(数字)
18-cons.conf.idx:消费者信心指数-月度指标(数字)
19-欧元银行同业拆借利率3M:欧元银行同业拆放利率3个月利率-每日指标(数字)
20-雇用人数:雇员人数-季度指标(数字)
输出变量(期望目标):
21-y-客户是否已认购定期存款?(二进制:“是”、“否”)
